From: Xavier Noria Date: 2006-04-17T07:13:26+09:00 Subject: Re: How extract data from a web site? On Apr 16, 2006, at 23:12, Ingo Weiss wrote: > I would like to use Ruby to read the content of a web site, and then > extract certain data from it. The site is machine generated so the > format doesnt' change, but unfortunately it is far from being valid > XHTML or similar. Some people already suggested HTML parsers. I have done a lot of crawling and wanted to add that if you need to extract data from a single page that is machine generated a simple regex is often enough. You need to make a choice depending on the real page and the kind of stuff to extract. -- fxn