From: pkellner Date: 2005-07-30T08:36:13+09:00 Subject: Re: confusion trying to get IMG tags from html page I was really hoping for some code or pseudo code. I'm new to ruby and have been thrashing over this for hours. I promise to put some back later when I know more about this. (and sadly, I'm not a regular expression wizard) Thanks Charles Steinman wrote: > pkellner wrote: > > I'm trying to download images from a web page that has them listed with > > html like what I've pasted below. Basically, I want to iterate through > > all the > I've tried a bunch of things with not much luck. Here is my last > > attempt. Any help would be appreciated. > > > > require 'net/http' > > require 'rexml/document' > > > > Net::HTTP.start('www.myphotowebsite.com') do |http| > > response = > > http.get('/terry/temp/2005-06-18%20Kiss%20of%20Death%203/') > > puts "Code = #{response.code}" > > puts "Message = #{response.message}" > > #puts "Body = #{response.body}" > > > > #parser = HTMLTree::XMLParser.new(false,false) > > #parser.feed(client.getContent(url)) > > xml=response.body > > > > xml.elements.each('//HREF]') do |node| > > > > end > > > > > > > > > > > > [IMG] > HREF="IMG_1516.jpg">IMG_1516.jpg 28-Jul-2005 08:59 > > 233k > > [IMG] > HREF="IMG_1517.jpg">IMG_1517.jpg 18-Jun-2005 08:03 > > 819k > > [IMG] > HREF="IMG_1518.jpg">IMG_1518.jpg 28-Jul-2005 09:00 > > 398k > > That isn't valid XML (tags without matching end-tags must have a > trailing slash), so the parser probably doesn't understand it. Assuming > the HTML isn't too complicated, you should be able to get the info with > regular expressions.