From: Mark Thomas Date: 2009-11-13T10:25:07+09:00 Subject: Re: Using Nokogiri On Nov 12, 11:52 am, jzakiya wrote: > On Nov 10, 10:29 pm, Mark Thomas wrote: > > > > > > OK, when I put Mark's code in a file and ran it (versus entering it in > > > a irb session) it DOES work. However, it doesn't capture the website > > > url, which 7stud's approach does. I haven't figure out how to do it > > > with this approach, and merely adding more items in xpaths doesn't > > > work. > > > > So Mark, how can your approach be used to capture the url add the end > > > of the data section? > > > > Here's the file I used with Mark's approach: > > > > File: scrape1.rb > > > ---------------------------- > > > require 'rubygems' > > > require 'open-uri' > > > require 'nokogiri' > > > > def scrape (id) > > > >   id = id.to_s > > >   url = "http://www.xyz.org/../../..ID=#{id}" > > >   doc = Nokogiri::HTML.parse(open(url)) > > > >   prefix = '//div[@class="sectionHeaderText"]/following-sibling::' > > >   xpaths = { > > >    :name => "#{prefix}b/text()", > > >    :addr => "#{prefix}text()[2]", > > >    :citystzip => "#{prefix}text()[3]", > > >    :country => "#{prefix}text()[4]", > > >    :phone => "#{prefix}text()[5]", > > >    :web => "#{prefix}text()[6]", > > >    :url => "#{prefix}text()[7]" > > > You'll need to modify that last line. Unlike the other items, the URL > > is not in a text node, it is the href attribute of the first > > element. So try: > > >     :url => "#{prefix}a[1]/@href" > > Yes, this allows me to capture the url I want (and sometimes ones I > don't want), and I'm able to post-process xpaths to get everything I > need. > >   xpaths = { >    :name => "#{prefix}b/text()", >    :addr => "#{prefix}text()[2]", >    :citystzip => "#{prefix}text()[3]", >    :country => "#{prefix}text()[4]", >    :phone => "#{prefix}text()[5]", >    :url => "#{prefix}a[1]/@href" >   } > > Now, I just need to understand completely WHY/HOW it works. :-) Let's take the first one as an example. I noticed that everything was after a div with the class "sectionHeaderText", so I started with that: //div[@class="sectionHeaderText"] The double slash is a wildcard that means the div can be anywhere. The part in brackets is called a predicate, and it constrains the expression. I like to think of it as a "such that" clause. So you can read the above as "a div such that the class is 'sectionHeaderText'." (Actually, it's the set of all divs for which it is true, so if you had multiple divs with the same class, it would return them all) Then I noticed that the items you wanted were not children of the div. The div closes before you get to the text you want. Even
tags are considered to be
which are self-closing. Therefore almost everything you want is at the same nesting depth, or in XPath terminology, they are siblings. The "following-sibling" is an XPath "axis" (see the W3C Schools XPath tutorial for details on these things). The name though was inside a element so I used the XPath expression to get the following sibling that happens to be a element: //div[@class="sectionHeaderText"]/following-sibling::b Then, how you get text from within a node is the XPath function text() which means all the text between tags, including whitespace. //div[@class="sectionHeaderText"]/following-sibling::b/text() And there you have the name. Now, the other things were text nodes between
elements. You could pull them all by asking for the set of text node siblings of the div: //div[@class="sectionHeaderText"]/following-sibling::text() But when you get more stuff than you want like that, you can index them like an array: //div[@class="sectionHeaderText"]/following-sibling::text()[2] and that happens to pull the street address. So hopefully you see how the XPaths were put together. Usually they are a bit simpler, but like 7stud said, it was pretty crappy HTML. -- Mark.