From: Mark Thomas Date: 2009-11-11T12:30:16+09:00 Subject: Re: Using Nokogiri > OK, when I put Mark's code in a file and ran it (versus entering it in > a irb session) it DOES work. However, it doesn't capture the website > url, which 7stud's approach does. I haven't figure out how to do it > with this approach, and merely adding more items in xpaths doesn't > work. > > So Mark, how can your approach be used to capture the url add the end > of the data section? > > Here's the file I used with Mark's approach: > > File: scrape1.rb > ---------------------------- > require 'rubygems' > require 'open-uri' > require 'nokogiri' > > def scrape (id) > >   id = id.to_s >   url = "http://www.xyz.org/../../..ID=#{id}" >   doc = Nokogiri::HTML.parse(open(url)) > >   prefix = '//div[@class="sectionHeaderText"]/following-sibling::' >   xpaths = { >    :name => "#{prefix}b/text()", >    :addr => "#{prefix}text()[2]", >    :citystzip => "#{prefix}text()[3]", >    :country => "#{prefix}text()[4]", >    :phone => "#{prefix}text()[5]", >    :web => "#{prefix}text()[6]", >    :url => "#{prefix}text()[7]" You'll need to modify that last line. Unlike the other items, the URL is not in a text node, it is the href attribute of the first element. So try: :url => "#{prefix}a[1]/@href"