From: James Edward Gray II Date: 2006-11-20T23:02:08+09:00 Subject: Re: Ruby screen scraping On Nov 20, 2006, at 2:45 AM, Paul Lutus wrote: > In this thread, the OP started out by examining the alternatives among > specialized libraries meant to address the general problem, but > apparently > never considered writing code to solve the problem directly. Starting out by looking for a library that does the hard work for you is a good first step, I would say. Do we really want to be discouraging that? > As to modern XHTML Web pages that can pass a validator, I know from > direct > recent experience that they yield to the simplest parser design, > and can be > relied on to produce a tree of organized content, stripped of tags and > XHTML-specific formatting, in a handful of lines of Ruby code. I've seen valid XHTML that wouldn't be much fun to parse. You still need to worry about whitespace, namespaces, the kind of quoting used, CDATA sections, ... James Edward Gray II