From: Bryan Murphy Date: 2002-06-20T05:45:26+09:00 Subject: Re: REXML in C > > >I have to deal with really large (300,000 line) XML files, and REXML, >well, just doesn't work for me. I really hate that since I really >like REXML. My idea, which I sort-of-started-but-haven't- >had-time-to-really-work-on-yet, was to back the REXML API with >libxml2. It has a fast parser and a fast XPath implementation, so I >think it could be very nice. I'm very interested in helping any >project that makes REXML faster. > >Steve > In all honesty, I really don't think you should be processing 300,000 line XML files with any DOM-like XML interface. There is a LOT of overhead and wasted memory in the DOM tree, and for an XML file that big it must be huge (not to even consider the time it takes to parse the XML into the DOM tree). Every time you access that information you're scanning the contents of the ENTIRE XML document in memory. With a document that size this is bad bad bad! What I think you really should be using is some sort of streaming parser or pull based parser when you are dealing with documents of this magnitude. You can build parsers that are considerably faster (orders of magnitude) and load data into much more compact (and applicable) data structures in memory. Yes, I know that writing a state based streaming parser is a bit harder than doing the same with REXML, but when you are dealing with this magnitude of data the tradeoffs are worth it in the long run imho (and building a good state based parser is a fun learning experience if you've never done it before)! Bryan