From: Clifford Heath Date: 2008-01-30T14:49:57+09:00 Subject: Re: Treetop parser (or PEG in general?) questions Phrogz wrote: > I suppose I'm suggesting that it isn't badly needed, since it's > trivial to add: > rule EOF > !. > end Oh, ok, I missed the . in your original. I thought that ! by itself was a special token. I agree, !. should work and there's no need to add it to the language. >> I think Nathan would oppose it, but I'd also like to add regex's >> as terminals for performance, so that a SyntaxNode isn't needed for >> every character. > I would be very, very in favor of this. Not just for performance, for > simplicity in consuming a few nodes that I don't need granularity on > and where it would be easier to match using a regexp. Certainly... and you'd need to save the MatchData so you can grab $1, $2, etc. I don't think it's even very hard... > I read a snippet that made it sound like Perl6 combines regexps and > PEG in some way; haven't looked into it any further to find out, > though. Not sure about PEG, but you can do recursive parsing and embed arbitrary code in Perl6 REs. > Hrm, not wild about implicit insertion into rules. (But then I'm just > a bumpkin.) No, I agree it's ugly, but it is essentially what other parser generators do by using a separate lexer. I just thought it would be cool to have a skip rule rather than some implicit one. It'd be a wart though. However, my original proposal still works and went down well with Nathan, which is to be able to declare a non-building rule using the keyword skip: skip white [ \t\r\n] end You still need to call that rule anywhere you want whitespace skipped, but you don't get nodes for it. As a further optimization, if you had (thing white)* then it could build an array of thing, instead of an array of sequences, each sequence containing one thing. Not sure how Treetop would know to do that though, as it only ever compiles a single rule at a time. > Additionally, something that I've wanted a few times (like the email > parser) would be a parser command that skips ALL node creation, Ok, but there needs to be something created to be able to memoize the parse. I'm not sure how much you'd save. BTW, for those who were hoping I might do all of this (or multi- language support) any time soon, well, I have another project underway that's taking all my time ATM :-). Clifford Heath.