From: Paul Lutus Date: 2006-11-21T02:15:05+09:00 Subject: Re: Ruby screen scraping Gabriele Marrone wrote: > > On 20/nov/06, at 03:50, Paul Lutus wrote: > >> array = page_content.scan(%r{

(.*?)

}m).flatten > > Please note that the P end tag isn't required in HTML 4.01: > http://www.w3.org/TR/html4/struct/text.html#h-9.3.1 Yes, I've just been converting all my site pages to XHTML, so I encountered this difference big-time. My solution made some assumptions, one being the OP's request -- >> My question is, how would i modify the >> code in order to get it to capture say a block of text such as: >> >> 

this is text that i want to scrape

>> >> any ideas? -- was based on his knowledge that the pages in fact contained paragraphs enclosed by

...

. The other assumption I made was based on context -- it seems the pages in question are machine-generated, so presumably can be relied on to have consistent syntax. -- Paul Lutus http://www.arachnoid.com