From: Bob Hutchison Date: 2008-01-14T21:44:44+09:00 Subject: Re: Stream Parsing with REXML Hi Robert, On 14-Jan-08, at 2:50 AM, Robert Klemme wrote: > Yep, looks pretty similar. To be honest I never used an XML pull > parser myself. Personally I prefer the push parser a bit because it > avoids the looping and decision logic based on event and element > type (in #work). Are there major advantages of pull parsers over > push parsers that I have overlooked so far? I use both and I'm comfortable with both. The combination of sax and pull gives you a much more complete toolkit for streaming XML. The biggest advantage (for me) to a pull parser is when you are parsing an XML file and using those events to drive some kind of algorithm. Either the parser or the algorithm is going to have to be able to suspend processing (and I mean save state by that) and resume when the other needs it to. This can be really hard to do for some algorithms, in the case of some it isn't possible without fundamentally changing the algorithm. If you read about pull parsers they will talk about where the flow of control (the 'outermost' loop) is in the program: in the parser or in the application. With sax the parser controls the flow, with pull parsers the application does. A subset of that situation is that the pull parser is just an object that you can pass around, and that can be *really* handy. This also means you can have more than one pull parser operating at a time, and that's incredibly confusing with sax. You can fake a pull parser by putting a sax parser in a different thread and streaming the events over some kind of channel with some kind of blocking protocol. Some implementations of pull parsers do exactly that, and that is *slow*, yet if you need it people will put up with it. If you've ever had the temptation to do that then you might want to think pull parser. It is pretty much trivial to write a sax parser based on a pull parser at very small cost (xampl-pp ships with an example that does that). So one trick that you can play is switching back and forth between sax and pull where appropriate. Junior programmers have a much easier time with pull parsers. It seems to provide a gentler route to sax based software. XML is a big enough step. There is some advantage to avoiding the need for a beginner to learn both XML and event based programming at the same time. Historically pull parsers tend to be very fast compared to sax parsers. This is not going to last, there's no technical reason for such a difference. In the Java world the speed difference is eroding for the good of all. Anyway, speed isn't the issue it used to be. If you poke around the web you'll still see a bunch of references to performance advantages, this doesn't apply anymore if you choose your parser correctly. Cheers, Bob ---- Bob Hutchison -- tumblelog at http://www.recursive.ca/so/ Recursive Design Inc. -- weblog at http://www.recursive.ca/hutch http://www.recursive.ca/ -- works on http://www.raconteur.info/cms-for-static-content/home/