From: ser@... (Sean Russell) Date: 2002-05-04T23:31:19+09:00 Subject: REXML and streaming APIs Hello, I'd like to solicit comments from the people in this group that use streaming XML parser APIs. I recently received an email from a user who needed REXML to do something it didn't, and he provided a change to REXML that enabled this behavior. Before I choose a direction REXML is going to go in with changes (additions, rather... I won't break backward compatibility) to the streaming API, I'd like to get feedback from people who actually use streaming APIs -- I don't use them, myself, so I don't feel confident about "laying down the law" about the best API. The questions I'd like answered are: Which will solve your problems best? Which seems the most intuitive? Which will involve the least code on behalf of the API user? Here are the current options: First, I'm going to add a StreamParser class. It would have the parse_stream method from Document. I can add methods to make it look more like a standard implementation of the listener pattern, where you can add and remove multiple listeners. IE, you'd get two new methods: add_listener( listener ) remove_listener( listener ) So stream parsing would change from: Document.parse_stream( stream, listener ) to p = StreamParser.new( stream, listener ) p.parse and you'd also be able to do: p = StreamParser.new( stream ) p.add_listener( listener1 ) p.add_listener( listener2 ) p.parse I'm also considering a step() method, which would be used to step through the document at your leisure: p = StreamParser.new( stream, listener ) step() # Parses the first tag encountered, and halts at the next step( true ) # Parses the next tag in its entirety step( "e" ) # Parses everything until it hits tag "e" I'm not very keen on this API, because (1) I doubt it would be very useful, and (2) it seems like a lot of work to use. What I like more is adding a more intelligent listener interface, similar to what NQXML provides: p = REXML::StreamParser.new(file) p.handle(:tag_start, %w(root level1 level2)) {|e| # do something with root, level1, and level2 tags } p.handle(:tag_start, %w(level3 level4)) { |e| # do something else with level3 and level4 tags p.handle(:text, %w(root level1 level2 level3)) {| e | # do something with the text nodes } p.parse and/or p = REXML::StreamParser.new(file) p.handle( listener, %w( root level1 level2 ) ) p.handle( listener, %w( level3 level4 ) ) p.parse The second looks like less code, but would be more because of the listener class definitions. However, I think they solve different problem domains; in the second case, you dedicate a class to handling all of the events within an element to a class, and in the first, you're only getting notified about a small domain of events. I can implement both, but I don't want to implement anything that isn't going to be used. One last thing. It seem petty of me to stick with "tag_start" and "tag_end" when SAX uses "start_element" and "end_element". At the cost of a little bloat, I can change the listener API to conform more to SAX and /mostly/ maintain backwards compatibility -- or I can throw out the old API, break backwards compatibility, and put in heafty warnings for the next few releases. Can I get a vote on this? Please CC me directly with your feedback, as I'm rarely in the newsgroup these days. Thanks! --- SER