From: jzakiya Date: 2009-11-09T03:35:06+09:00 Subject: Using Nokogiri I'm trying to scrape some data off websites using nokogiri require 'rubygems' require 'open-uri' require 'nokogiri' #using the latest 1.4.0 url = 'http://www.whateverwebsitenameis.org' doc = Nokogiri::HTML(open(url)) This gets me data off the website I want to scrape. The segment of the site I want looks like this (from FF 'view source' ) -------------------------------------------------------------------------

Association Detail

DETAIL DIRECTORY RESULTS
1) Some Institute name

2) some address
city, st zip
3) 4) United States
5) 6) Phone: 7) 8) (123) 456-7890
9) 10
11) Web address: www.xyz.org


Back to Search Results

Search Again --------------------------------------------------------------------------------- I want to scrap and collect the data between lines 1-11, ie, name, address, city, st, zip, United States, phone number, and line 11 I want the website url: 'http://www.xyz.org' I can find the beginning of this section of code by doing this: doc.css('h2').each do |elem| puts elem.content end which displays 'Association Detail' I am having problems using this as the starting point to parse the data in lines 1-11 which contain the specific 'Association Detail' details. I've tried it with 'xpath' and 'search' according to the example here: http://rdoc.info/projects/tenderlove/nokogiri but there's something I'm just not getting correctly when I use other elements get info from. My system is Windows XP, Ruby 1.8.6, Nokogiri 1.4.0 Thanks in advance for any help.