From: William James Date: 2006-03-08T10:48:40+09:00 Subject: Re: Is there link extractor or similar html processing libs for Ruby Desireco wrote: > Hi, > > in cool Perl there are a bunch of libraries that process html files and > > help you when you need to extract info. I remember hearing something > for Ruby as well, if someone had experience with this, it would help me > > if he could point me in right direction. Basically I need to extract > links and info from html pages. > > Thanks. > > > Zeljko Dakic > http://www.dakic.com class String def xtag(s) result = [] scan( %r! < #{s} (?: \s+ ( [^>]* ) )? / > | < #{s} (?: \s+ ( [^>]* ) )? > ( .*? ) !mix ) \ { |unpaired, attr, data| h = { } ( unpaired || attr || "" ). scan( %r{ ( \w+ ) \s* = \s* (?: ( ["'] ) ( .*? ) \2 | ( \S+ ) ) }x ) { |k,q,v,v2| h[k.downcase] = (v || v2) } block_given? ? ( yield [ h, data ] ) : result << [ h, data ] } result end end DATA.read.xtag('a'){|atr,txt| puts atr['href'], txt } __END__ foo bar foo bar upcoming HTML 3.2 reference. All the is A, with the attribute HREF. Help |