From: Nat Pryce Date: 2002-05-07T20:50:07+09:00 Subject: Re: REXML and streaming APIs I'd like to see a "pull" parser, rather than a push parser for a number of reasons. 1) It much easier to process XML data in a recursive descent style with a pull parser compared to a push parser. 2) It is easier to convert code that recurses over a document tree into code that parses a stream using recursive descent, than it is to convert code that recurses over a document tree into code that parses a stream using a push parser. This would be useful when code written to process small documents has to be modified to work with larger documents. 3) It is easy to write an adapter class that converts a a pull parser into a push parser with an adapter class, while it is much more complex to convert a push parser into a pull parser. I think the best API would be similar to a lexer. Client code would instantiate an object that returned a stream of tokens describing the XML document: begin element, end element, attribute, pcdata, cdata, entity, etc. Cheers, Nat. ________________________________ Dr. Nathaniel Pryce B13media Ltd. Studio 3a, Aberdeen Business Centre, 22/24 Highbury Grove, London, N5 2EA http://www.b13media.com ----- Original Message ----- From: "Sean Russell" Newsgroups: comp.lang.ruby To: "ruby-talk ML" Sent: Saturday, May 04, 2002 3:31 PM Subject: REXML and streaming APIs > Hello, > > I'd like to solicit comments from the people in this group that use > streaming XML parser APIs. > > I recently received an email from a user who needed REXML to do > something it didn't, and he provided a change to REXML that enabled > this behavior. Before I choose a direction REXML is going to go in > with changes (additions, rather... I won't break backward > compatibility) to the streaming API, I'd like to get feedback from > people who actually use streaming APIs -- I don't use them, myself, so > I don't feel confident about "laying down the law" about the best API. > > The questions I'd like answered are: > > Which will solve your problems best? Which seems the most intuitive? > Which will involve the least code on behalf of the API user? > > Here are the current options: > > First, I'm going to add a StreamParser class. It would have the > parse_stream method from Document. I can add methods to make it look > more like a standard implementation of the listener pattern, where you > can add and remove multiple listeners. IE, you'd get two new methods: > > add_listener( listener ) > remove_listener( listener ) > > So stream parsing would change from: > > Document.parse_stream( stream, listener ) > > to > > p = StreamParser.new( stream, listener ) > p.parse > > and you'd also be able to do: > > p = StreamParser.new( stream ) > p.add_listener( listener1 ) > p.add_listener( listener2 ) > p.parse > > I'm also considering a step() method, which would be used to step > through the document at your leisure: > > p = StreamParser.new( stream, listener ) > step() # Parses the first tag encountered, and halts at the next > step( true ) # Parses the next tag in its entirety > step( "e" ) # Parses everything until it hits tag "e" > > I'm not very keen on this API, because (1) I doubt it would be very > useful, and (2) it seems like a lot of work to use. > > What I like more is adding a more intelligent listener interface, > similar to what NQXML provides: > > p = REXML::StreamParser.new(file) > p.handle(:tag_start, %w(root level1 level2)) {|e| > # do something with root, level1, and level2 tags > } > p.handle(:tag_start, %w(level3 level4)) { |e| > # do something else with level3 and level4 tags > p.handle(:text, %w(root level1 level2 level3)) {| e | > # do something with the text nodes > } > p.parse > > and/or > > p = REXML::StreamParser.new(file) > p.handle( listener, %w( root level1 level2 ) ) > p.handle( listener, %w( level3 level4 ) ) > p.parse > > The second looks like less code, but would be more because of the > listener class definitions. However, I think they solve different > problem domains; in the second case, you dedicate a class to handling > all of the events within an element to a class, and in the first, > you're only getting notified about a small domain of events. I can > implement both, but I don't want to implement anything that isn't > going to be used. > > One last thing. It seem petty of me to stick with "tag_start" and > "tag_end" when SAX uses "start_element" and "end_element". At the > cost of a little bloat, I can change the listener API to conform more > to SAX and /mostly/ maintain backwards compatibility -- or I can throw > out the old API, break backwards compatibility, and put in heafty > warnings for the next few releases. Can I get a vote on this? > > Please CC me directly with your feedback, as I'm rarely in the > newsgroup these days. > > Thanks! > > --- SER