From: Erika Date: 2008-12-10T17:15:01+09:00 Subject: Re: pop3 body email --0-215819000-1228897304=:41638 Content-Type: text/plain; charset=us-ascii Hi, I tried using Hpricot in the following way: ..... Connect with POP3 and get the email ..... email = TMail::Mail.parse(m.pop) body2=Hpricot(email.body) elements = body2.search("/html/body/table//td").collect{|k| k.inner_html.split(',') unless k.inner_html =~ / Info1Info2
Info1Info2
Info1Info2
Info1Info2
Info1Info2
Info1Info2
So the general rule is that i have 3 tables and I need to check that for Info1 the correct Info2 is shown. Every Info1 / Info2 can vary. Is there a better way to have in to arrays for example Info1 and Info2. Because my solution is ok, only I subtract all the info from the html code which I need to parse one more time. Thanks, Erika ________________________________ From: Robert Dober To: ruby-talk ML Sent: Tuesday, December 9, 2008 11:02:39 PM Subject: Re: pop3 body email On Tue, Dec 9, 2008 at 9:29 PM, Brian Candler wrote: > Erika wrote: >> Hi, >> >> I have to check the body of the email which is a html code, like the >> following: >> Text1: >> >> Text2 >> >> I have to check that for "Text1" what "Text2" is shown, which can vary >> for different reasons. > > If you are already able to read the E-mail via POP3, and the body > consists of a single text/html part, then this just becomes a question > about parsing HTML. Hpricot is usually cited as the best library for > doing that. I'd say it's not worth digging about with regexps when you > can do the job properly. Thx a lot there, but please do not underestimate regexen, especially as they will become more powerful in 1.9. Are you aware of dependencies of Hpricot notwithstanding that it is a wonderful tool. Being an old Unix guy however I feel that you do not need a full fledged library + dependencies if a three liner can do the job. However if this was only an example and OP needs more parsing Hpricot is a very sensible way to go. Cheers R --0-215819000-1228897304=:41638--