From: Phil Tomson Date: 2008-01-30T14:47:15+09:00 Subject: Re: Treetop parser (or PEG in general?) questions On 1/28/08, Clifford Heath wrote: > Phil Tomson wrote: > > This is good to know. So you could also do this, correct?: > > > > rule inline_atom > > ( numz { def bar; ... end} / wrrd { def bar; ... end} ) { def foo; ... end} > > end > > > > so that each alternative gets it's own bar method. > > Yes. > > > So how would one go about extracting one valid syntactic element (the > > port_decl in this case) from surrounding elements that one doesn't > > care about? > > You're trying to skip any amount of stuff up to the port declaration, > then parse that, then skip to the end. Now first I'll ignore that your > "stuff" can presumably contain comments, which may contain the word > "port", but you do it like this: > > rule vhdl_file_wth_port > ( !'port' . )* port_decl .* > end > > This says to parse any number of single characters as long as you aren't > looking at the word "port", then parse the port_decl and skip the rest. OK, this approach seems to work (I'll need to try the additions you oulined later as well to prevent matching words which have "port" in them). Just to up the ante a bit: In VHDL entity declarations can have port declarations in them as can component declarations: entity Foo is port( a,b : in bit; c : out bit); end Foo; --later component CPU is port( clock : in bit; data_bus : inout bit_vector( 15 downto 0); address_bus : out bit_vector( 31 downto 0) ); end CPU; There should only be 1 entity declaration in a file, but there could be multiple component declarations (or none) in a file. I'd like to extract the entity port declaration and component port declarations (keeping a list of component ports). Any suggestions? Again, there's lot of syntax that can occur around these things and a I don't care about any of that. I just want to get the connections between things. > rule string > "'" ( '\\' [befntr\\'] / .)* "'" > end > > rule white > [ \t\n\r]+ > end BTW: I've noticed that \w doesnt' seem to work as the univerasl whitespace designator. I tried this comment rule: rule comment '--' [0-9a-zA-Z\w]* [\n] end But it didnt' work when I had a comment like: -- this is a comment But this one did: --thisisacomment Then I changed the rule to add a ' ': rule comment '--' [0-9a-zA-Z ]* [\n] end And that worked. Phil