From: Nat Pryce Date: 2002-06-05T01:07:37+09:00 Subject: Re: HTML Parser suggestions wanted I think implementing your HTML parser as an add-on to REXML is a great idea. If your parser builds a REXML document, you will inherit any improvements to REXML, such as faster XPath support, with no extra effort on your part. Also it will be easier for people writing applications that process both XML and HTML if they only need to learn one document API. The only difficulty I can see is getting the document to write itself out as HTML rather than XML. Cheers, Nat. ________________________________ Dr. Nathaniel Pryce B13media Ltd. Studio 3a, Aberdeen Business Centre, 22/24 Highbury Grove, London, N5 2EA http://www.b13media.com ----- Original Message ----- From: "Ned Konz" To: "ruby-talk ML" Sent: Tuesday, June 04, 2002 4:30 PM Subject: HTML Parser suggestions wanted I've written an HTML parser that builds trees from HTML source. After I wrote it, I discovered REXML, which does the same thing for XML. Then I made an add-on that uses REXML's XPath support to do XPath queries on the resultant HTML tree. In my test version, these queries return REXML tree elements, rather than my HTML tree elements. Having two very similar tree structures (HTML and REXML) smells, to me. The fact that they have somewhat different APIs confuses even me. What I'm wondering (and would like your input on): 1. Should I just require REXML and not bother with my own tree elements? I could, after all, just build a REXML document instead. This has the disadvantage of requiring yet another package to be installed, though. 2. If I don't build an REXML tree, I could still return my own elements from XPath queries. That is, I could use REXML transparently and not expose the user to any of REXML's elements. Would this be a preferable way to provide XPath support? Thanks, -- Ned Konz http://bike-nomad.com/ruby/ GPG key ID: BEEA7EFE