From: windwiny Date: 2013-05-12T18:13:28+09:00 Subject: Re: parse both string and url using Nokogiri xpath --047d7b6dcdd2de12ce04dc817e76 Content-Type: text/plain; charset=UTF-8 Mayby using a temp variable ? links = doc.xpath('//div/a[@href]') links.map do |x| [x.text,x['href']] end => [["reference", " http://example.com"]] 2013/5/12 Soichi Ishida > ruby 1.9.3 > nokogiri 1.5.5 > > Say, a web page has a link, > > reference > > I would like to get both the url and text, "http://example.com" and > "reference". > > First, access to the page that contains this link. > > doc = Nokogiri::HTML(open(url)) > > then, > > name = doc.xpath('//div.../a').text > url = doc.xpath('//div.../a/@href).text > > It works. But the problem is this is parsing twice separately. > If you want to apply the same procedure to many links that exist in a > single page, it seems inefficient. > > Is there anyway to produce both url and text by single parse? like > > def parse_link_and_text (xpath) > ... > end > > p parse_link_and_text('//div...') > > gives a hash > > => {'reference' => 'http://example.com'} > > ? > > -- > Posted via http://www.ruby-forum.com/. > > --047d7b6dcdd2de12ce04dc817e76 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable
Mayby using a temp variable ?

=C2= =A0 =C2=A0 links=C2=A0=3D doc.xpath('//div/a[@href]')
=C2=A0 =C2=A0 links.= map = do |x| [x.text,x['href']] end =C2=A0 =C2=A0=3D> [["referenc= e", "http://example.com"]= ]


2013/5/= 12 Soichi Ishida <lists@ruby-forum.com>
ruby 1.9.3
nokogiri 1.5.5

Say, a web page has a link,

=C2=A0 <a href=3D"= http://example.com">reference</a>

I would like to get both the url and text, "http://example.com" and
"reference".

First, access to the page that contains this link.

doc =3D Nokogiri::HTML(open(url))

then,

name =3D doc.xpath('//div.../a').text
url =3D doc.xpath('//div.../a/@href).text

It works. =C2=A0But the problem is this is parsing twice separately.
If you want to apply the same procedure to many links that exist in a
single page, it seems inefficient.

Is there anyway to produce both url and text by single parse? like

def parse_link_and_text (xpath)
=C2=A0 =C2=A0...
end

p parse_link_and_text('//div...')

gives a hash

=3D> {'reference' =3D> 'http://example.com'}

?

--
Posted via http://= www.ruby-forum.com/.


--047d7b6dcdd2de12ce04dc817e76--