From: Mark Probert Date: 2005-03-01T03:40:36+09:00 Subject: Re: Parsers vs. Homemade "Parsing" via REs Hi .. On Sunday 27 February 2005 06:54, Randy Kramer wrote: > I have the need to translate several megabytes of TWiki marked up text to > HTML. ... > > * If I did create the proper grammar rules, would parsing using > something like YACC be faster than a bunch of RE replacements? > I think that you mileage may vary, depending on the nature of the text markup and how complicated you want the HTML. If it is something like tag replacement, then regexp is probably the way to go. It will be, on the whole, faster to implement for such a "one-off" job. > * Any recommendations for a parser in Ruby? I think there are a couple, > I've been doing some Googling / reading and have come across references to > parse.rb and (iirc) something called Coco (??). > I am responsible for one of the Coco/R implementations (coco/rb). The "something" is a what's called an attributed LL(1) parser-generator, as opposed to YACC, which is LALR. These are just different ways that the parser attacks decomposition of the target file. You can find out more at http://www.scifac.ru.ac.za/compilers/ There are some basic tradeoffs. Coco is very fast and the grammar and scanner are all in one place (YACC uses LEX as a tokeniser, which is a separate program and input file). YACC is somewhat more flexible in grammars it can produce. For example, it is not possible to write a Ruby parser in COCO/r, though you can do it YACC. However, for 99% of little-languages, that is not such a big problem, just avoid constructs like 'goto'. Regards, -- -mark. (probertm at acm dot org)