From: Gregory Brown Date: 2006-12-18T01:22:31+09:00 Subject: Re: Grabbing data off a webpage On 12/17/06, Bil Kleb wrote: > OK, so I haven't done this in years. > > What's the "modern" way of grabbing the data off > a webpage, e.g., > > http://yorkcountyschools.org/mves/arlist/3-3.4.htm > > My initial attempt has been focused on Hpricot, > > require 'rubygems' > require 'open-uri' > require 'hpricot' > doc = Hpricot(open('http://yorkcountyschools.org/mves/arlist/3-3.4.htm')) > > and I can find doc/"th" and doc/"tr", but what's > the best way to cram them into an array of structs > or something? I've actually been needing to do something like this for work and haven't gotten around to it, so I'll take a stab at it. require "ruport" column_names = (doc/"th")[1..-1].map { |r| (r/"p").text } rows = (doc/"tr")[3..-1] parsed_rows = rows.inject { |s,a| s << (a/"td").map { |r| (r/"td").text } } table = parsed_rows.to_table(column_names) Now, I've pastied some of the things you can do from here, because they wont translate to email well. http://pastie.caboo.se/28169 Note, my hpricot code is sort-of hackish, cleaning that up might be a good idea, but Ruport[0] might still be a good idea for representing the data. Hope this helps! -greg [0] http://ruport.infogami.com