From: Robert Klemme Date: 2013-03-27T17:25:00+09:00 Subject: Re: Nokogiri help parsing HTML On Wed, Mar 27, 2013 at 12:05 AM, Jesús Gabriel y Galán wrote: > On Tue, Mar 26, 2013 at 11:40 PM, Paul Mena wrote: >> My goal is to strip out everything between the "address" and "pre" tags >> and to output only: > > OK, so you want every tag that is a sibling of address and pre and is > within those two. I have found this StackOverFlow answer: > > http://stackoverflow.com/a/6555376 > > which applied to your problem: > > 1.9.2p290 :001 > require 'nokogiri' > => true > 1.9.2p290 :002 > s = < 1.9.2p290 :003"> > 1.9.2p290 :004">
> [...snip...] > 1.9.2p290 :031 > doc = Nokogiri::HTML(s) > 1.9.2p290 :039 > > doc.xpath("//address/following-sibling::node()[count(.| > //pre/preceding-sibling::node())=count(//pre/preceding-sibling::node())]") > => [#, > # children=[#, > #, > #, > #, > #, > #]>, > #] > > will return a node set that contains the required nodes. Your version also outputs

tags, doesn't it? A modified version of yours irb(main):036:0> dom.xpath('//address/following-sibling::*//text()').each {|n| p n} # # # # # => 0 irb(main):037:0> dom.xpath('//address/following-sibling::p//text()').each {|n| p n} # # # # => 0 Here's another approach: find everything under

but not under
: irb(main):032:0> dom.xpath('//div[@class="mail"]//text()[not(ancestor::address)]').each {|n| p n} # # # # # # # => 0 TIMTOWTDI :-) Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/