From: Michael Morin Date: 2008-09-05T03:36:14+09:00 Subject: Re: Extract data from email - Tmail, Hpricot George Cooper wrote: > Hi all, > > I have an html email that I would like to parse. > > The problem I'm having is removing all html tags and getting past the > header information. Then I want to extract all the information per row > to put into a database. > > the email is pasted here: http://pastie.textmate.org/265259 > > I have tried Tmail, but can't seem to extract just the body. Then I > tried Hpricot and wasn't sure what to use before the .inner_html. So > basically I'm very lost on where to start. > > Any help is appreciated. > > Thanks! It would help if you posted some code that didn't work, so people can have a better idea of what you're trying to do. Tmail should have been able to parse that without problem, however, extracting the body is easy. The box follows the empty line. You could use something like split, but duping such huge strings could be slow. When you read the mail, try to read a line at a time until you get the empty line, then read the rest into a buffer for hpricot. -- Michael Morin Guide to Ruby http://ruby.about.com/ Become an About.com Guide: beaguide.about.com About.com is part of the New York Times Company