From: 7stud -- Date: 2009-11-09T08:25:56+09:00 Subject: Re: Using Nokogiri jzakiya wrote: > I'm trying to scrape some data off websites using nokogiri > > require 'rubygems' > require 'open-uri' > require 'nokogiri' #using the latest 1.4.0 > > > url = 'http://www.whateverwebsitenameis.org' > > doc = Nokogiri::HTML(open(url)) > > This gets me data off the website I want to scrape. > > The segment of the site I want looks like this (from FF 'view > source' ) > > ------------------------------------------------------------------------- >

Association Detail

> >
DETAIL > DIRECTORY RESULTS
> > 1) Some Institute name

> 2) some address
city, st zip
> 3) > 4) United States
> 5) > 6) Phone: > 7) > 8) (123) 456-7890
> 9) > 10
> 11) Web address: target="_Blank">www.xyz.org
> >

> > Back to Search Results a>

> > >
Search Again > > > --------------------------------------------------------------------------------- > > I want to scrap and collect the data between lines 1-11, ie, name, > address, city, st, zip, United States, phone number, and line 11 I > want the website url: 'http://www.xyz.org' > > I can find the beginning of this section of code by doing this: > > doc.css('h2').each do |elem| puts elem.content end > which displays 'Association Detail' > > I am having problems using this as the starting point to parse the > data in lines 1-11 which contain the specific 'Association Detail' > details. I've tried it with 'xpath' and 'search' according to the > example here: http://rdoc.info/projects/tenderlove/nokogiri > > but there's something I'm just not getting correctly when I use other > elements get info from. > > My system is Windows XP, Ruby 1.8.6, Nokogiri 1.4.0 > > Thanks in advance for any help. You aren't really searching by css, which would involve things like searching for tags based on their 'class' attribute or 'id' attribute. Because the

tag doesn't have any attributes, you are simply searching by tag name, so you could do this instead: doc.xpath('//h2').each do |h2| puts h2.content end That uses xpath notation to find all h2 tags on the page. Then you might write something like this: doc = Nokogiri::HTML.parse(html) doc.xpath('//h2').each do |h2| if h2.content == "Association Detail" puts "---" puts h2.next.content puts "---" end end Knowing you can do that will enable you to write something like this: results = [] doc.xpath('//h2').each do |h2| if h2.content == "Association Detail" curr_elmt = h2 while (curr_elmt = curr_elmt.next) curr_content = curr_elmt.content results << curr_content break if curr_content.include?("Web address:") end end end results.each do |result| puts "--start--" puts result puts "--end--" puts end output= --start-- DETAIL DIRECTORY RESULTS --end-- --start-- Some Institute name --end-- --start-- --end-- --start-- --end-- --start-- some address city, st zip United States Phone: (123) 456-7890 ) Web address: www.xyz.orgBack to Search Results a>Search Again --end-- As you can see, the html is pretty bad, so your results aren't that great. You will have to figure out how to extract the data you need from those strings. -- Posted via http://www.ruby-forum.com/.