From: Eric Mahurin Date: 2008-02-05T01:46:33+09:00 Subject: Re: [QUIZ] Parsing JSON (#155) ------=_Part_7398_18487196.1202143599602 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: 7bit Content-Disposition: inline On Feb 3, 2008 11:45 PM, Clifford Heath wrote: > Eric Mahurin wrote: > > I'm biased, > > First up, it sounds like your package is a significant achievement, > and I don't want anything here to imply I'm not commending you on > it. But different approaches suit different goals... Thanks! > but I don't think Treetop is super-clean and pure. Take a look > > at my rubyforge grammar package if you want something really clean and > pure: > > :-) To me, to use a separate grammar is *more pure* than to use > a Ruby DSL. I can see how you would argue the opposite however... > Your example parsers are much harder to read than Treetop however, > compare: > space = Grammar::Element[Set[?\ ,?\t].duck!(:==,:include?)].discard > with > skip space > [ \t] > end > I agree. It is ugly. The ugliness can be easily fixed though: 1. when making a parser, derive from Grammar so that namespace is available 2. E = Element, so that you can have a shorthand for the most common atomic Grammar 3. Performance-wise, I found that using a Set for matching was slower than just having a few (ranged) alternatives. 4. Grammar::Element.[] is simply an alias for Grammar::Element.new. My plan is that in the next release, I'll make all the classes "callable" by defining methods of their name that call new. This is better when .new can take a block (like Recurse). 5. I made a bad choice of #== for the matching method for the pattern argument. #=== would be better and is what I'm using in my dev code (and the Grammar0 I posted). This would allow it to work out of the box with Ranges too. So, you could do this: alpha = E(?a..?z) | E(?A..?Z) So, doing 1..3, in the released grammar v0.5, you'd have this instead: space = (E[?\s] | E[?\t]).discard (ok, so TT doesn't have a skip keyword yet, soon...) > That's what I mean by clean and pure anyhow. Not pure Ruby, > but pure grammar. > > Nathan's concept is that "grammar" should become a Ruby keyword, > and should introduce the grammar meta-grammar, rather than using > existing Ruby syntax. I think that polyglot approach is much better > than the DSL approach, since the Ruby grammar isn't well-suited to > many important languages. I don't think ruby needs any new keywords. It already has more than it needs in my opinion. There is enough power with just classes, methods, blocks, and operator overloading. The point is, if/when Ruby uses a parser whose grammar is composable > with a DSL grammar you define, the result is truly seamless and > the syntax is unrestrained. Existing DSL syntax imposes too many > restraints. A language with a composable grammar would be truly > excellent! > > > * specify language grammars in a BNF-like ruby DSL (no new grammar > language > > to deal with!) > > I think I've argued why that is a bad thing :-), though it definitely > has good things about it also - like your filters etc... > > Ruby DSL is far too awkward to express CQL - which is my primary > project at present. Don't know anything about CQL to answer. I would like to make what I have general enough. > - integration of lexer/parser with other code is seamless since > everything > > is ruby > > Nice, but it won't help me when I need to target C# or Java. > Most important grammars need to be compilable to more than > one target language. Treetop's not there yet, but it's easy > to see how to do it, whereas it won't ever happen with yours > AFAICS. It definitely could happen with mine. I do ruby code generation now, but could have another Grammar-like class (or another mechanism) to generate code for another language. I would have to add a new method for dumping out the generated code (right now it is only eval'ed immediately). In my dev code, I'm also using ruby2cext to compile it on the fly. I'll try to have a general language infrastructure in the next release. > - (sub)grammars are objects and you combine them using operators and > > methods > > Can't get much easier than just "include otherGrammar", which > works in Treetop (mix the generated modules together). > > > * complete API can be described in only a little more than a hundred > lines > > Similarly to Treetop, AFAICS. > > > * infinite backtracking (not by default because of performance and > > functionality, you specify what needs it) > > Does this memoize, or give exponential performance? Or is it optional? In general, the parsers it generates are LL(1). It makes decisions one character/token at a time. But, if you use the #lookahead method on any grammar, it will give a new grammar that handles the case when the grammar fails in the middle (backtracks to the position where the grammar started). For most languages I've seen this is rarely needed, but it is in your back pocket to be used when needed. I haven't done anything to try to optimize it as it isn't needed that often. > * can make any kind of pipeline of lexers/parsers and multi-thread them. > > * even with its pure ruby/no-regexp lexers/parsers, it can go > head-to-head > > against Regexp-based stuff. The flexibility of Regexp is limited, so I > > don't see the point since 'grammar' gets enough performance. > > * don't need to read the file into a string to parse it > > * and more > > Very cool stuff. > > > With 'grammar' you have complete control of the parsing result. > Without > > actions, the output stream is a copy of the input stream. You can > group > > things and discard things efficiently. One extreme would be to > interpret > > the input on the fly and not build anything. > > Yes, this is excellent, and something I've always wanted from a parser > generator. You should add your json parsers to the examples in the gem! Yes, next release. ------=_Part_7398_18487196.1202143599602--