From: Sean Russell Date: 2001-11-14T08:52:01+09:00 Subject: [ruby-talk:25106] Re: REXML behavior change RFC John Carter wrote: > In another way I worry about large documents. It would mean that the API > itself specifies that the engine has to slurp the whole document before it > can give you the first element to handle. Well, REXML *does* parse the entire document before giving you the first element to handle. This is the only (easy) way to provide XPath support. The other option, which I've considered, is to have REXML parse the document in a separate thread, and have some sort of mechanism for hanging XPath queries until the document is parsed. Please note that to support XPath, either (1) the entire document has to be in memory, (2) the document must be parsed multiple times, or (3) the document is parsed once and written do disk in some serialized format, which is then read back in a chunk at a time when queries are performed. I've implemented (3) in Java before, and it involves quite a bit of work. Perhaps I'll look into it someday, after REXML is a fairly "clean" toolkit, and I'm not busy supporting XML conformance as much. (2) is just not an option. Therefore, we have (1). > With XML being targeted as the way unrelated databases exchange data, we > can look to see some very large streams of data going by... If you want to use Streams, you can use the streaming parser of REXML, which *doesn't* keep the whole document in memory. On the other hand, streaming parsing doesn't support XPath, so it doesn't have to. -- --- SER