From: Ryan Davis Date: 2009-12-05T07:52:49+09:00 Subject: Re: ruby language parser in ruby On Dec 4, 2009, at 13:45 , Brian Candler wrote: > Ryan Davis wrote: >> I think your best bet at this time is ruby_parser. > > Question for you: as far as I can make out, the handling of embeeded > linebreaks such as > > a = b + > c > > is handled via state kept in the lexer. Can I ask why you did it this > way? My default answer to nearly any of these types of questions is going to be "because that's how MRI does it". In this case, yes, the lexer and parser are keeping a shared variable called lex_state that knows whether it is in the middle of an expression (and many other states) so here the trailing '+' keeps the expression open. lex_state needs to die. It basically only exists because the language is a tangled mess and it was designed with LR parsing in mind (AFAICT, lex_state is a symptom of not really knowing where you are contextually, because you're parsing bottom up). I'd like to not have lex_state and many of the complications that come with it. I'm not sure if it'll still be ruby at that point, but I'm giving it a go to see. > I am thinking it ought to be possible to do this in the grammar, > e.g. > > expr: expr '+' opt_nl expr > | expr '-' opt_nl expr > > opt_nl: > | nl you're going to have those EVERYWHERE... but yeah, it should be possible. Write good tests from the beginning. > Also: can you summarise how you're handling expressions nested within > string literals, e.g. "abc #{foo} def" ? There is no way to summarize that. It is horrible and I hate it, but I'm not in a position to make it work better with the current architecture.