From: "Mauricio Fernández" Date: 2004-01-23T06:08:37+09:00 Subject: Re: BibTeX parser On Fri, Jan 23, 2004 at 03:34:27AM +0900, Daniel Carrera wrote: > Hi all, > > I have a difficult problem and I need some smart people to give me a hand. > So I knew where to go for that. :) > > I'm trying to figure out how to write a parser fo BibTeX files. [...] > If I only had to deal with {brackets} I could use an n-ary tree. And to > parse it, I would start with one node, move one character at a time. > Every time I see a { I'd make a new node. Every time I saw a } I would > come back up. > > > Now, when you and "double" quotes, the problem becomes more complicated, > but doable. I could first extract all the quotes and use an array where > quoted and non-qutoed text alternates (for instance) and then parse using > the brackets to make an n-ary tree. > > > But if I have 'single' quotes also, things can get very complicated. [...] > And at this point I am truly at a loss. Looks like you're doing the parser by hand... wouldn't it be easier with, say, racc? As for the lexer, you could simply split (well, not String#split but you get the idea) on spaces & special chars ({}"'\); creating a grammar to handle this should be fairly easy. Another advantage is that you could build an AST and use it to represent the data. If needed you could simplify it later to transform "recursive" nodes (i.e. those resulting from recursive productions) into arrays; this is more convenient and IIRC it's what you'd get with Rockit. You might also want to try the latter, but in my past experience I found it to be too buggy :-( mmm I guess Coco/Rb could be a good option too, since you also get a lexer, and LL(1) should be enough for this. -- _ _ | |__ __ _| |_ ___ _ __ ___ __ _ _ __ | '_ \ / _` | __/ __| '_ ` _ \ / _` | '_ \ | |_) | (_| | |_\__ \ | | | | | (_| | | | | |_.__/ \__,_|\__|___/_| |_| |_|\__,_|_| |_| Running Debian GNU/Linux Sid (unstable) batsman dot geo at yahoo dot com If loving linux is wrong, I dont wanna be right. -- Topic for #LinuxGER