From: Jimmy Thrasher Date: 2001-09-30T03:00:02+09:00 Subject: [ruby-talk:21827] Re: ANN: HtmlDomGenerator 0.1 > # You should have received a copy of the GNU General Public License > # along with this program; if not, write to the Free Software > # Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA > > No GPL to be found :-) Hrm. I thought there was a line in there that says, "but if you didn't, here's where to look." Oops. :) > 1. It's not clear (to me) what you mean by "a DOM". Are you planning > to implement the DOM API? Or more of a home-grown tree-based > interface? (In which case you'd be best off avoiding the term "DOM".) I suppose I could just call it a tree, but it looks like the DOMs I've used. ;) However, I was planning in future releases to make it jibe with a real DOM API. I whipped this up yesterday and today so I could parse HTML files much more easily than otherwise. > 2. If this is a general SGML and/or XML tool, I wouldn't use "HTML" in > the name. Agreed. The only thing is, I don't do any DTD checking, so consider this a non-validating parser that was convenient at the time for what I want to do. I shouldn't call it anything, by that token. Or, GenericMarkupLanguageTreeStructureGenerator. :) > 3. "Well-formed HTML" is sort of a contradiction in terms, or at least > redundant. HTML has a DTD; therefore, HTML documents are either valid > (and, incidentally, well-formed) or invalid. If a document is > well-formed but not valid HTML, then it isn't HTML. (I guess there's > such a thing these days as "well-formed [DTD-less] SGML", though I > think that only exists as a way to make sure that all XML documents > are also SGML documents.) Right, my stuff is currently a quick hack to get a tree structure I can deal with. It may improve, if I don't end up using something else. > 4. I'm not sure what "handling" interleaved tags would entail, other > than dying with a fatal error :-) Ask Netscape or IE. :) Or HtmlTidy. > 5. You might get some mileage out of existing Ruby tools in this area, > including Jim Menard's NQXML, which has a tree parser. Would NQXML handle HTML
tags? I need something that will handle HTML. Shall I update the description and names to communicate the lack of DOM-iness, and lack of validation? Or does it matter? Jimmy