From: Robert Klemme Date: 2008-01-14T16:50:56+09:00 Subject: Re: Stream Parsing with REXML On 13.01.2008 20:54, Bob Hutchison wrote: > Hi Marc, Robert: > > On 13-Jan-08, at 1:04 PM, Robert Klemme wrote: > >> I have attached another example of an idiom that I use frequently. >> This may not be needed in your case but who knows? Basic idea is >> that the event listener creates nested listeners (one per element) >> and hands processing off to them while maintaining a stack of >> listeners. That way you can do different processing based on >> element name, attributes or whatever. Might be overkill for your >> simple example but OTOH if you do need to do complex processing >> steps based on elements this might be exactly what you need. For >> example, you can store any information you need in a nested listener >> and do all the processing in end tag. > > That idiom is really frequent in the stuff that I do with XML. It also > happens to be a case where a pull parser can make things clearer. Pull > parsers are event based like SAX parsers are, they do not build up > trees in memory, and they are very fast. The code below is almost > equivalent to yours Robert, as far as I could make it anyway. The > parents method is different, I considered that as well but wanted to be able to traverse parent nodes top down for the example. After all it was just a quick hack. :-) For more robustness and library fitness I would certainly change a few things. > and I habitually add a node level > corresponding to the document since there are things (comments, > processing instructions, etc) that are allowed outside of the root > element of an XML file. If you look closely that happens in my version as well (see StreamListener#initialize and #tag_start). > I wrote that xampl-pp parser in May 2002. It is on source forge > but not as a gem. Hmm, maybe I should do something about that, I've > made some fixes to very obscure problems over the years, but never > released anything. Anyway the point isn't xampl-pp so much as what a > pull parser looks like. Many people don't know about them and if they > do they often don't know why they are useful. I think the libxml2 > library does pull parsing (and that'll be very fast) and I think REXML > does too (it has a few problems with namespaces last I looked). > Anyway... > > All the interesting stuff happens in the work method. > > #!/bin/env ruby > > require "xampl-pp" > > class Thing > > attr_accessor :tag_name > attr_accessor :parent > attr_accessor :text > > def initialize(parent, name) > @parent = parent > @tag_name = name > @text = "" > end > > def Thing.start(xml) > pp = Xampl_PP.new > pp.input = xml > > Thing.new(nil, 'doc').work(pp) > end > > def parents > p = self > while p > yield p > p = p.parent > end > end > > def work(pp) > parents do |par| > print par.tag_name, " " > end > puts > > while not pp.endDocument? do > case pp.nextEvent > when Xampl_PP::START_ELEMENT > Thing.new(self, pp.name).work(pp) > when Xampl_PP::END_ELEMENT > print " ", text.inspect, "\n" > return > when Xampl_PP::TEXT || Xampl_PP::CDATA_SECTION || > Xampl_PP::ENTITY_REF then > text << pp.text > end > end > end > end > > Thing.start(DATA) Yep, looks pretty similar. To be honest I never used an XML pull parser myself. Personally I prefer the push parser a bit because it avoids the looping and decision logic based on event and element type (in #work). Are there major advantages of pull parsers over push parsers that I have overlooked so far? Kind regards robert