From: Luiz Vitor Martinez Cardoso Date: 2008-01-06T11:32:39+09:00 Subject: Re: html parser with regex, how to solve? ------=_Part_47050_21970385.1199586762823 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: 7bit Content-Disposition: inline Thanks I do it! Regards, Luiz Vitor Martinez Cardoso. On Jan 6, 2008 12:15 AM, Joe wrote: > try this: > > ele.text.sub('$', '') > > Joe > > On Jan 5, 2008 8:34 PM, Luiz Vitor Martinez Cardoso > wrote: > > Thanks much! This really works ;) > > > > Now i`m having a new problem (very simple), the output is $1999,00, how > i > > can remove a $? I will need convert this to a float number ;) > > > > Regards, > > Luiz Vitor Martinez Cardoso. > > > > > > On Jan 5, 2008 11:10 PM, s.ross wrote: > > > > > Don't use the regex. Let hpricot do what it's good at: > > > > > > $ irb > > > >> require 'rubygems' > > > >> require 'hpricot' > > > >> html = '

149.00

' > > > >> doc = Hpricot.parse(html) > > > >> ele = doc.search('h3.zmp') > > > >> puts ele.text > > > => 149.00 > > > > > > In your code, your @content will be searchable the same way. Hpricot > > > will give you a collection of all h3's with class 'zmp'. > > > > > > http://code.whytheluckystiff.net/doc/hpricot/ > > > > > > Hope this helps. > > > > > > > > > On Jan 5, 2008, at 4:07 PM, Luiz Vitor Martinez Cardoso wrote: > > > > > > > Yeah, > > > > > > > > I`m trying to develop a simple application using ruby (when this > > > > works i > > > > will pass to rails). I need get the source code from a URL, and find > > > > for > > > > this string: > > > > > > > >

$299.99

> > > > > > > > wow, but i need search for not only 149.00, but for all possible > > > > numbers, my > > > > friend suggest this: > > > > > > > >

*$\d+\.\d{2}.*

> > > > > > > > i think this works! but i need other thing... look my code: > > > > > > > > #!/usr/bin/ruby > > > > > > > > require 'hpricot' > > > > require 'open-uri' > > > > > > > > @content = Hpricot(open(" > > > > http://www.newegg.com/Product/Product.aspx?Item=N82E16855101066")) > > > > > > > > now how i can find for

*$\d+\.\d{2}.*

? > > > > > > > > @content.search("

*$\d+\.\d{2}.*

") is broken ;( > > > > > > > > how i can solved this? > > > > > > > > > > > > thanks for you attention, > > > > Luiz Vitor Martinez Cardoso. > > > > > > > > > > > > > > > > -- > > > > Regards, > > > > Luiz Vitor Martinez Cardoso [Grabber]. > > > > (11) 8187-8662 > > > > > > > > rubz.org - engineer student at maua.br > > > > > > > > > > > > > > > -- > > Regards, > > Luiz Vitor Martinez Cardoso [Grabber]. > > (11) 8187-8662 > > > > rubz.org - engineer student at maua.br > > > > -- Regards, Luiz Vitor Martinez Cardoso [Grabber]. (11) 8187-8662 rubz.org - engineer student at maua.br ------=_Part_47050_21970385.1199586762823--