From: Alex LeDonne Date: 2007-11-20T03:55:51+09:00 Subject: Re: Scraping from a website On Nov 19, 2007 1:45 PM, cskilbeck wrote: > Hi, > > I need to extract everything between
and
on a website > (there's only one table on the page. So far I have: > > require 'open-uri' > page = open('http://xxx.html').read > page.gsub!(/\n/,"") > page.gsub!(/\r/,"") > inner = page.scan(%r{.*(.*).*}m) Untested, but try: inner = page.scan(%r{.*]*>(.*).*}m) > print inner > > but inner is empty - any ideas? > > If I substitute line 2 with > > page = '123456
789 > > I get inner = 456, which is correct. If you try page = '123
456
789', it will fail again. You only want to capture up to the next closing angle bracket. What's happening is that the second .* is matching the contents of the entire table, up to the closing angle bracket of the last tag (probably ) right before the , and inner gets only the leftover whitespace inbetween. So only capture characters that are NOT a closing angle bracket. -Alex