From: Clifford Heath Date: 2008-01-28T17:24:57+09:00 Subject: Re: Treetop parser (or PEG in general?) questions Phrogz wrote: > I've been looking for something like treetop for a while now. Very > excited to have found it, and to play with it. So was I - thanks Nathan! - and many of the recent improvements are mine. > Let's assume that I'm trying to parse the following (wiki) markup: ... snip... > Three questions (in particular, of many) jump out at me: > LINE ANCHORING > In the above, there is 'block' content, and 'inline' content. Headers, > paragraphs, and tables are block level items. Among other things, this > means that their markup is only valid at the start of a line. For > example, a line that started " = Hi =" would not constitute a valid > header. In Regexp land, this would be handled simply with a ^ anchor. > How do you handle this in Treetop? Treetop (and PEGs in general) requires no separate lexer - all lexing can be done with the same algorithmic efficiency as you get from a typical DFA-based lexer anyhow. As a result, Treetop has no notion of whitespace or newlines - they're just characters, and you have to have syntax rules that match them. If you want a rule to match after a newline, call it after a newline, or at the start of the text). If you want a rule to match only if followed by a newline, follow it with: & '\n' You can use any rule after & - the rule must match but the input won't be consumed by the calling rule. > HANDLING EOF > As seen in the example above, a paragraph (or any block level element, > really) is allowed to not have a newline if it's the last thing in the > file. Do you handle this case normally by just preprocessing the input > and shoving a newline on the end if it doesn't exist, That's what I'd do. > or is there a > way in Treetop to recognize the /\Z/ anchor from a Regexp? No. There's no EOF symbol. Perhaps there should be.. > BACKREFERENCES > A valid heading must match this regexp: /^(?:=+) (.+) \1$/ > It must start at the front of the line with one or more = > characters. > It can have anything (including some = characters). > It ends with the same number of = characters, which must be followed > by a newline. I don't think there's a way of saying that the trailing =s must be equal in number to those matched by the leading '='+. In general, there's no way to inject custom code that affects the parsing process (what ANTLR calls a semantic predicate), which is a pity. It's something I want sometimes too, so if you can suggest a clean enough way to specify it, I'll think about implementing it. > Other questions I won't dive into: what's a reasonable way to eat > inline content (words) while allowing inline markup? Your problem here if I read it correctly is that you want to read anything that's not markup. Here's where the ! operator comes in. You might have a rule called "markup", which matches any markup, and a rule "word" that you call like this: rule words (!markup word)* end That will match any sequence of zero or more "word", where no word matches "markup". Here, the rule following ! may be of any complexity, as with &. This is essentially the "C comment matching problem". Here's what I use for C-style comments, C++ style comments, and whitespace: rule s # Optional space S? end rule S # Mandatory space (white / comment_to_eol / comment_c_style)+ end rule white [ \t\n\r]+ end rule comment_to_eol '//' (!"\n" .)+ end rule comment_c_style '/*' (!'*/' . )* '*/' end > Am I necessarily > going to end up with a tree for a paragraph that has one child for > each word (or letter)? You'll have one leaf per leaf rule (lexical rule) - but you don't need to look at it, you can use "text_value" of any node, which is just the substring of the input spanned by that rule. I hope that's some help. My CQL parser might give you some more ideas, at . It's a different style of language than what you're parsing, but also needs large amounts of backtracking at times. Clifford Heath.