From: Eric Mahurin Date: 2005-11-06T09:16:03+09:00 Subject: Re: parser performance comparisions ------=_Part_43698_7580707.1131236160605 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Content-Disposition: inline On 11/5/05, Steven Jenkins wrote: > > Eric Mahurin wrote: > > The racc memory usage could also be helped if you don't > > generate the entire token stream in memory first. I'm not sure why ther= e > are > > no racc examples with low memory usage (don't read the entire file into > a > > string and don't generate the entire token stream in memory up front). > > Probably because for most applications it doesn't matter. But the > example I sent you doesn't read the entire file into a string. It does > load the entire token stream into an array, but that could easily be > changed by running the lexer in a separate thread and using a SizedQueue > for the token stream. Maybe a five-line change. > > For this appliation, I really don't care about memory consumption. The > largest file I'll ever parse is 100 times smaller than the physical > memory on the machine, much less the available virtual memory. > > Steve > > Do you even need a separate thread? Why can't you just have #next_token jus= t parse the next token and return it? From what I saw of your lexer, it looke= d like that would be possible, but then again I have almost no racc experienc= e (I've only learned enough to try to benchmark against it). When you go mutli-threaded, you do have to make sure your doing enough work in each thread because thread switching is relatively expensive. For my multi-threaded test, increasing my token buffer from no buffering (lexer an= d parser synchronized) to a 300 token buffer made a couple orders of magnitud= e difference in performance. To me, if a parser is reading in a whole file into memory and/or generating all the tokens of the file in memory up front, it doesn't seem like a "real= " parser. Regexp fits in this camp (parses from a string - in memory). With some of features that are being added to regexes (look at some of perl's advanced/beta features where you can embed code and do recursion), you coul= d call a regex a parser, but it doesn't seem like one to me because it needs the whole thing it is parsing in memory. It just seems strange to me that the common racc usage does things this way. You could actually optimize a parser quite a bit if you knew everything it was parsing was in memory (regexes make this assumption for doing backtracking - similar to lookahead for a standard parser). ------=_Part_43698_7580707.1131236160605--