From: Miroslaw Niegowski Date: 2009-01-27T17:38:57+09:00 Subject: Re: Page crawling and URL grabbing 2009/1/27 Patrick L. : > Hey guys, > I'm trying to write an application that goes onto a website (istockphoto > specifically), opens up istockphoto.com/file_browse.php and grabs the > URLs of the photos that appear there. > > It's my first time doing something like this. I'm reading some > documentation right now...but a hand would be greatly appreciated. I'm > not really sure how to do regex on an html file...or even find the right > stuff within that file. I'm guessing its.. > > open('http://www.istockphoto.com/file_browse.php/') do |f| > f.find # dot something something > end Try Mechanize. It's easy : agent = WWW::Mechanize.new agent.user_agent_alias='Mac Safari' page = agent.get('http://www.istockphoto.com/file_browse.php'); page.links.text(/jpg/) ...