From: James Britt Date: 2006-05-12T05:31:13+09:00 Subject: Re: extract DIV from HTML Ryan Leavengood wrote: > You want to use XPath. Check out the XPath section of the following > document: > > http://www.germane-software.com/software/rexml/docs/tutorial.html > > For your specific case, the XPath query would be something like this: > > //div[@class='whatever'] Depends. If the large HTML file is indeed a LARGE file, then REXML XPath may be CPU/memory intensive. If so, the REXML pull or stream parsers work quite nicely for this sort of thing. I'd also suggest WWW::Mechanize. (I also saw a suggestion for RubyfulSoup, but my experience is that it bogs down on large files. ) -- James Britt http://www.ruby-doc.org - Ruby Help & Documentation http://www.artima.com/rubycs/ - The Journal By & For Rubyists http://www.rubystuff.com - The Ruby Store for Ruby Stuff http://www.jamesbritt.com - Playing with Better Toys http://www.30secondrule.com - Building Better Tools