From: David Vallner Date: 2007-02-19T20:39:47+09:00 Subject: Re: Screen-scraping Hmm. On Mon, 19 Feb 2007 12:01:01 +0100, Peter Szinek wrote: > If you can do without javascript for a moment, Having since dared the HTML, it seems at least for some of the apps, not really. > 1) Mechanize and Hpricot are super great in themselves I never actually used either, I can only vaguely guess at the scope - Hpricot doing the low-level parsing and cleanup, Mechanize the higher-level data extraction from the result of that. > 3) I am planning to invest a lot of time into scRUBYt! - I am just > releasing the next version as I write this mail, my TODO list has about > 200+ items and the community seems to be very active, so I got already > tons bug reports, feat requests and even patches (and the whole thing is > out for about 2 weeks) Woo, new anal retention sink? (COWER BRIEF MORTALS.) Who knows, I might even get around to actually fixing other peoples' bugs in my spare time when I'm not suppressing homicidal tendencies from doing so at worktime. (Although I expect my spare hacking time will be spent coding said screenscrapers in the nearby future.) > amazon_stuff = Scrubyt::Extractor.define do > > fetch 'http://www.amazon.com' > fill_textfield 'field-keywords', 'logitech keyboard' > choose_option 'url', 'Computers & PC Hardware' > submit I like this API. > > stuff do Where'd the stuff variable come from? > item_name "Logitech diNovo Edge ( 967685-0403 )" I love and adore my dNE too ;P > price "$169.98" > end > end > > amazon_stuff.to_xml.write($stdout, 1) > Scrubyt::ResultDumper.print_statistics(amazon_stuff) Right, I suppose it goes on the List of Things To Try on the saner of the webapps. And after that Excel automation for the paperwork done -that- way (unsurprisingly the most laborious of them all.) David Vallner