From: "Florian Groß" Date: 2005-04-25T22:37:01+09:00 Subject: Re: announcing RubyLexer 0.6.0 vikkous wrote: >>Which is exactly what I thought would be a good way of >>extending. This looks good. > > Everything may not be as simple as this one case was. The fact that the > first example you gave turned out to be pretty easy is encouraging, but > I think we're likely to run into something really nasty before you are > happy. Hm, that ought to be not too much of a problem. I'm okay with having a look at some of the internals for that kind of things. >>It would then be very nice if I could lex until I see the 'if' then >>say 'give me an atomic expression' which would parse until >>the 'then' and then say 'give me an atomic expression' again >>which would parse until the 'end'. Basically I don't want to >>match paired things (parentheses, do .. end, class definitions >>etc.) at the transformation level. > > In general, 'get the next expression' is a problem that requires a > parser, not a lexer. Have you looked at ParseTree? Of course you have. > > In this case however, you are in luck. Delimited expressions, that > start and end with ( and ), or begin and end, or whatever, are already > discovered by my lexer. (During the development of RubyLexer, I > discovered that it had to be half-a-parser as well, in order to > correctly get all the information that's needed to lex correctly.) The > information you want is already being gathered by RubyLexer, it's just > not available in a public interface. We should negotiate such an > interface since you seem to need it. What you propose, 'get the next > expression', is not one I want to do. RubyLexer does not deal in > abstractions larger than tokens... at least, not on a public level. I > am, however, willing to emit 'advisory' tokens at certain points in the > token stream, (several such types of tokens are being emitted already) > which should allow you to do what we want, if we design it carefully. Hm, I am not sure if that is enough for this case. The condition part of a if or something else will after all not always be surrounded by ( and ) or begin and end or something similar. Advisory tokens (which would tell me that I am now entering the condition of if and now leaving it and now entering the action part of it and so on) might do this. However, you are right in that this is not usually the task of a lexer. In the past I have frequently had trouble with the distinction of lexing and parsing in real language parsing -- most languages require you to keep some context for actually tokenizing them. Ruby, for example, requires that your lexer knows about all kinds of quoted Strings and where they end and interpolated expressions inside them. I'm not sure of where to best draw the line so it's probably better to let you decide. > On the other hand.... the reason I chose not to emit advisory tokens > for this particular case is that the complimentary tool to RubyLexer is > intended to be Reg, which can find nested pairs of braces and the like > pretty easily. Have you looked at Reg at all? I realize that I only > released it yesterday, and as of yet it's only half-working because > critical features are as yet unimplemented, but I think it might be > just the thing for the types of preprocessors you have in mind. Heh, I didn't realize that you were also the author of that library so I did not draw the connection. I have, however, marked those two threads as something I will have to examine. (They are now colored red.) I'm watching Reg with growing interest -- I'm not sure if I have already told this to you (I remember telling the author of "BNF-like grammar specified DIRECTLY in Ruby"), but I have also done something vaguely similar -- I have done an object-oriented way of constructing and combining Regular Expressions. What you have done is something better. I'm especially interested in how the LALR parser, Reg and RubyLexer might all work together. Any way of getting some sample code? I'm aware of the fact that this is all subject to change as long as you have not implemented all the necessary features like look-ahead, but getting a quick overview would still be nice. > Reg might not be able to easily tell 'if' the postfix operator from > 'if' the value in current RubyLexer output. Since one requires an end > and the other doesn't, that can be troublesome to deal with. 'do' is > also a pain, now that I think of it. All these cases are handled > correctly in RubyLexer, we just have to find an appropriate > (token-based, not expression-based) interface. I would be pretty much okay with the advisory tokens idea -- it sounds like meta-tokens that tell me about the context. >>For doing code transformations it is of course also important that >>you can turn back the stream of tokens into a String easily. I did >>this with IRB's lexer by using the .line_no and .pos methods of >>tokens, but that was not too good a match, actually. > > So what would be a good match? I don't see why this should be a > problem. My implementation of Token implements to_s, which returns the > ruby code corresponding to the token; ususally, this is exactly the > same as the code that created the token originally. There's also a > offset method, which returns the position of the token in the input > stream, relative to the very beginning. Tokens don't have a #line_no, > but you can get the same information from FileAndLineTokens. This does sound good. Having an offset ought to actually be better than separate character and line numbers as well. >>I did this with IRB's lexer by using the .line_no and .pos >>methods of tokens, but that was not too good a match, actually. > > Wait,,,, so you wrote irb's lexer? One of my wishlist items is to > integrate RubyLexer with irb among others.... how hard do you think > this will be? Nope, not really. I've just used it out of IRB. Integrating it ought to be possible, but I'm not sure why that would be necessary. > Well, it was unexpected for me. Much to my embarassment; I thought I > was an expert at this. I must say many elements of this got me very > confused at first, and obviously I never put all the pieces together. > Congratulations. > > Ps: I haven't figured out why this breaks RubyLexer yet, but I will. Good luck. :) > Pps: putting tricky stuff in eval strings and the like won't break the > lexer (yet). To the lexer, it's just a string. Yup, same for IRB.