From: Wyss Clemens Date: 2002-01-15T23:50:15+09:00 Subject: RE: grabbing stuff from web pages what an honor to help Ron Jeffries ;-) b) just this morning I *hacked* (i.e. to be refactored ;-)) this working code snippet: require 'net/http' def getJpeg(webserver, urlPath) print "getting #{urlPath}\n" h = Net::HTTP.new(webserver, 80) begin resp, data = h.get(urlPath, nil ) rescue print " FAILED!\n" return end # writing the jpeg jpegFilename = urlPath.sub(/.*\/(.*?)$/, '\1') File.open(jpegFilename,"w") { |fh| fh.syswrite(data) } end ... getJpeg should work for you, too. Regards Clemens > -----Original Message----- > From: Ron Jeffries [mailto:ronjeffries@REMOVEacm.org] > Sent: Dienstag, 15. Januar 2002 15:33 > To: ruby-talk@ruby-lang.org > Subject: grabbing stuff from web pages > > > Part of my web site has recommended books. I use the cover jpegs from > Amazon as part of my page. The rule with Amazon is that you have to > download the picture and put it on your site, not link to theirs. > > So I was starting out with Dave and Andy's example: > > require 'net/http' > > h = Net::HTTP.new('www.xprogramming.com', 80) > resp, data = h.get('/index.htm', nil) > if resp.message == "OK" > puts data > data.scan(/ end > > figuring that I could then come up with some way to get and save the > image ... I remember doing it once before as an experiment, though I > can't find the code now. > > But when you set up to hook to the amazon book page, which can be > accessed this way in a web browser > > http://www.amazon.com/exec/obidos/ASIN/0201708426/ > > so that the code looks like this: > > require 'net/http' > > h = Net::HTTP.new('www.amazon.com', 80) > resp, data = h.get('/exec/obidos/ASIN/0201708426/', nil) > if resp.message == "OK" > puts data > data.scan(/ end > > You get an error 302. I suppose that it's all that magic redirection > that they're using to get you to the real page, but I don't > know enough > to be sure. > > So I'd appreciate help on > > a) getting the right page address given that I only know the number > that comes after ASIN, and > b) saving the jpeg once I get its src address. > > Thanks! > > Ronald E Jeffries > http://www.XProgramming.com > http://www.objectmentor.com > I'm giving the best advice I have. You get to decide whether > it's true for you. >