From: Benoit Cerrina Date: 2001-11-18T21:48:30+09:00 Subject: [ruby-talk:25750] Re: Why not xmlparser? (was: Re: XML support in the standard lib;whatexactly?) > XMLParser isn't much faster than the native Ruby XML parsers, except in > Stream parsing, and it is much slower in many other common XML tree > operations. I don't see any reason to take on the burden of maintaining > the expat sources if we're not going to see a significant speed increase. > I haven't investigated the memory usage of XMLParser compared to NQXML or > REXML, so that may be a factor. You just said it was significantly faster when parsing as a stream. In my experience working with xml in c++ and java, stream parsing is almost always the way to go and tree parsing is almost always the way people start. It usually works that way: you have a parser which support DOM+SAX a developper needs to do something which usually end up being reading and writing xml which represent a persistent view of an internal tree structure. The developper looks at both apis, thinks SAX is too complex (which its not but the problem is that its event based and many developper think its too complex, I think that pull apis like MS parser in .NET or xpp http://www.extreme.indiana.edu/soap/xpp/ have all the advantages of event based SAX like api while being easier for the average developper to wrap his mind around) then he looks at the DOM, sees a tree says here I know what it is. He tries to use it, it is long and painfull since the DOM is an ugly api with a lot of traps to fall into. Then the project is done, it is slow it takes 10 times more memory than necessary (due to the huge dom memory consumption plus having both the native structures and the dom structures in memory). Then some other guy has to rewrite everything with the sax, the implementation ends up taking lot less code and going a lot faster. Please don't discasrd the speed of stream parsing since in fact it is the only one really relevant. Also do not discard the idea of pull parsing. Finally to me there is one time where a tree api is usefull it is when writing the xml. When parsin (i.e. reading) it is usually always a bad choice. Benoit