From: Peter Szinek Date: 2006-11-20T07:33:50+09:00 Subject: Re: Ruby screen scraping Chris Gallagher wrote: > OK that code all works great but i have one last question :) > > This is allowing me to scrape the values of the class values on tags and > any other attribues such as that. My question is, how would i modify the > code in order to get it to capture say a block of text such as: > >

this is text that i want to scrape

Hmm this is hard to tell just by this example. If you need ALL

s, then those can be queried by this XPath: //p I am not sure what are you using now, but in Hpricot this would be: doc = Hpricot(open("http://stuff.com/")) results = doc/"//p" If you are still using, HTree, query this XPath there for the same results. However, I guess you want something more sophisticated than ALL the

s. Well this is where the troubles begin with screen scraping: you need to figure out some rules which extract *exactly* what you want - usually it is not that hard to come up with rules that extract more or less, but much worse to find the right ones... To solve this problem, you need to tell us what do you want - i.e. an example page, and a set of objects you would like to extract. Cheers, Peter __ http://www.rubyrailways.com