From: Robert Klemme Date: 2008-01-14T22:17:24+09:00 Subject: Re: Stream Parsing with REXML 2008/1/14, Bob Hutchison : > On 14-Jan-08, at 2:50 AM, Robert Klemme wrote: > > > Yep, looks pretty similar. To be honest I never used an XML pull > > parser myself. Personally I prefer the push parser a bit because it > > avoids the looping and decision logic based on event and element > > type (in #work). Are there major advantages of pull parsers over > > push parsers that I have overlooked so far? > > I use both and I'm comfortable with both. The combination of sax and > pull gives you a much more complete toolkit for streaming XML. I read that in the light of your following remarks. Other than that I do not see a difference with regard to the XML parsed, i.e. both approaches are equally powerful. > The biggest advantage (for me) to a pull parser is when you are > parsing an XML file and using those events to drive some kind of > algorithm. Either the parser or the algorithm is going to have to be > able to suspend processing (and I mean save state by that) and resume > when the other needs it to. This can be really hard to do for some > algorithms, in the case of some it isn't possible without > fundamentally changing the algorithm. If you read about pull parsers > they will talk about where the flow of control (the 'outermost' loop) > is in the program: in the parser or in the application. With sax the > parser controls the flow, with pull parsers the application does. I see, that's definitively a situation I was not aware of. I guess I just haven't been in a situation where I needed to employ such algorithms. And yes, the instance in control (parser vs. application code) seems to be the biggest difference. I can see how this makes a difference for some algorithms. > A > subset of that situation is that the pull parser is just an object > that you can pass around, and that can be *really* handy. This also > means you can have more than one pull parser operating at a time, and > that's incredibly confusing with sax. :-) > You can fake a pull parser by putting a sax parser in a different > thread and streaming the events over some kind of channel with some > kind of blocking protocol. Some implementations of pull parsers do > exactly that, and that is *slow*, yet if you need it people will put > up with it. If you've ever had the temptation to do that then you > might want to think pull parser. > > It is pretty much trivial to write a sax parser based on a pull parser > at very small cost (xampl-pp ships with an example that does that). So > one trick that you can play is switching back and forth between sax > and pull where appropriate. Yep. Basically a push parser can be implemented as thin wrapper around a pull parser. > Junior programmers have a much easier time with pull parsers. It seems > to provide a gentler route to sax based software. XML is a big enough > step. There is some advantage to avoiding the need for a beginner to > learn both XML and event based programming at the same time. I haven't thought of that either but sounds like a valid point. > Historically pull parsers tend to be very fast compared to sax > parsers. This is not going to last, there's no technical reason for > such a difference. In the Java world the speed difference is eroding > for the good of all. Anyway, speed isn't the issue it used to be. If > you poke around the web you'll still see a bunch of references to > performance advantages, this doesn't apply anymore if you choose your > parser correctly. Bob, thanks for taking the time and put together this elaborate explanation! Kind regards robert -- use.inject do |as, often| as.you_can - without end