From: vikkous Date: 2005-04-27T02:54:30+09:00 Subject: Re: announcing RubyLexer 0.6.0 > would produce a token stream similar to > > # pardon me if my way of representing this is not at all compatible > # with RubyLexer's design -- I need to get familiar with it soon > [KeyWord['if'], IfConditionStart, VariableOrMethod['condition'], > IfConditionEnd, KeyWord['then'], IfActionStart, > VariableOrMethod['action'], IfActionEnd, KeyWord['end']] > > And I think that that would be easier to analyze than the non- > annotated token stream. Of course you would still have to do > nesting counting to be able to extract the sections, but I think > that would be reasonable for simplicity's sake. Ok, fair enough. Maybe this way is easier after all. > Hm, why is that? Could it not use the rules it uses for parsing > for one-token-at-a-time-ahead lexing? I just can't see this. The lexer rules' input is the source file, but the parser's is the parse stack -- which comes from the lexer's output ultimately.... this can be a very powerful way to compose pattern matchers, but in the end different rule sets are used with 2 different inputs. The lexer and parser can run interleaved, and the lexer can get information from the parser to help interpret things (this is sometimes called "cheating", but it isn't; it's often the easiest way). But there's still the two rule sets. I don't know if it's possible to have 1 rule set do both at once, but the idea is intruiging. > I'm not sure whether not having lexing and parsing more > unified has benefits or downsides with your approach. I > guess I will just have to write a Joy interpreter using all this. > Do you think that that can already be done or is there > features missing that would make it wise to delay this further? I took a little look a joy. Hoo-boy. I'm guessing this language is pretty easy to parse. I would say reg is not ready for anything significant until it has backreferences and substitutions. At that point, it's got match-and-replace, and retrieval of arbitrary match subexpressions. If you think you can live without those, I'd say go for it. There are some problems with the backtracking engine, but so far as I can see, only a whole lot of ambiguity causes the problems, so it's _probably_ ok for most things. > Had I wanted to make this yet more difficult I would have > mixed multiple styles of newlines. ;) > > Now I actually do wonder if using CRLF instead of LF does > anything special to newline-delimited literals on any > platforms. Sure enough, I translated to unix format and the problems disappeared. Using a dos newline as a delimiter in a fancy string is just a little difficult for me because I had always assumed string delimiters were a single character... hrm. Here documents need this functionality to really support dos newlines correctly too, I think.