From: Gregory Brown Date: 2006-12-18T01:25:55+09:00 Subject: Re: Grabbing data off a webpage On 12/17/06, Gregory Brown wrote: > On 12/17/06, Bil Kleb wrote: > > OK, so I haven't done this in years. > > > > What's the "modern" way of grabbing the data off > > a webpage, e.g., > > > > http://yorkcountyschools.org/mves/arlist/3-3.4.htm > > > > My initial attempt has been focused on Hpricot, > > > > require 'rubygems' > > require 'open-uri' > > require 'hpricot' > > doc = Hpricot(open('http://yorkcountyschools.org/mves/arlist/3-3.4.htm')) > > > > and I can find doc/"th" and doc/"tr", but what's > > the best way to cram them into an array of structs > > or something? > > I've actually been needing to do something like this for work and > haven't gotten around to it, so I'll take a stab at it. > > require "ruport" > column_names = (doc/"th")[1..-1].map { |r| (r/"p").text } > rows = (doc/"tr")[3..-1] > parsed_rows = rows.inject { |s,a| > s << (a/"td").map { |r| (r/"td").text } > } > table = parsed_rows.to_table(column_names) > > Now, I've pastied some of the things you can do from here, because > they wont translate to email well. > > http://pastie.caboo.se/28169 Yuck, seems to have made a mess of the text output. Here it is better formatted: http://pastie.caboo.se/28170/text