From: james@... Date: 2002-06-05T00:50:54+09:00 Subject: RE: HTML Parser suggestions wanted > > I've written an HTML parser that builds trees from HTML source. After > I wrote it, I discovered REXML, which does the same thing for XML. I started looking at your parser (thanks!) and wanted to load the resulting HTML into REXML. > > Then I made an add-on that uses REXML's XPath support to do XPath > queries on the resultant HTML tree. In my test version, these queries > return REXML tree elements, rather than my HTML tree elements. .. and then I wanted to do XPath queries. So far, so good! > > Having two very similar tree structures (HTML and REXML) smells, to > me. The fact that they have somewhat different APIs confuses even me. > > What I'm wondering (and would like your input on): > > 1. Should I just require REXML and not bother with my own tree > elements? I could, after all, just build a REXML document instead. > This has the disadvantage of requiring yet another package to be > installed, though. True, but REXML is part of the Windows PragProg install, and it's quick and easy to install from source, anyway. > > 2. If I don't build an REXML tree, I could still return my own > elements from XPath queries. That is, I could use REXML transparently > and not expose the user to any of REXML's elements. Would this be a > preferable way to provide XPath support? Wouldn't REXML still need to be installed? BTW, quick question: Is there any documentation (aside form the terse API docs) about your parser? Is there a way to grab the parsed HTML and assign it to a string? All I've seen is how to write to a file. Thanks, James > > Thanks, > -- > Ned Konz > http://bike-nomad.com/ruby/ > GPG key ID: BEEA7EFE > >