From: "Jesús Gabriel y Galán" Date: 2013-03-27T08:05:17+09:00 Subject: Re: Nokogiri help parsing HTML On Tue, Mar 26, 2013 at 11:40 PM, Paul Mena wrote: > I'm relatively new to Ruby (and therefore Nokogiri) and am trying to > parse some HTML that will ultimately be written to a MySQL database. In > the interim, I'm writing it to a text file for troubleshooting purposes. > > Here's the relevant piece of the HTML I'd like to parse: > > >
>
> > From: Paul David Mena < href="mailto:pauldavidmena_at_gmail.com?Subject=Re:%20twilight">pauldavidmena_at_gmail.com> >
> Date: Tue, 26 Mar 2013 18:13:21 > -0400
>
>

> Line 1 >
> Line 2 >
> Line 3 >
>

> --
> Paul David Mena
> --------------------
> pauldavidmena_at_gmail.com
> 
> Received on Tue Mar 26 2013 - 22:13:23 > EDT >
> > > My goal is to strip out everything between the "address" and "pre" tags > and to output only: OK, so you want every tag that is a sibling of address and pre and is within those two. I have found this StackOverFlow answer: http://stackoverflow.com/a/6555376 which applied to your problem: 1.9.2p290 :001 > require 'nokogiri' => true 1.9.2p290 :002 > s = < 1.9.2p290 :004">
[...snip...] 1.9.2p290 :031 > doc = Nokogiri::HTML(s) 1.9.2p290 :039 > doc.xpath("//address/following-sibling::node()[count(.| //pre/preceding-sibling::node())=count(//pre/preceding-sibling::node())]") => [#, #, #, #, #, #, #]>, #] will return a node set that contains the required nodes. Hope this helps, Jesus.