From: Esmail Bonakdarian Date: 2006-09-30T11:45:05+09:00 Subject: Re: newbie: how to find & extract a string from a file Hi, Thank you so much for these pointers. Am I correct in assuming that REXML comes as part of standard Ruby? If so I think I will go that route first. I could cobble something together using various Linux tools (grep and awk come to mind), but I want something in Ruby (because I want to learn it) and also because it will be more portable, for instance to the XP platform. I appreciate you taking the time to post this and the references. If you have any other ideas/approaches, I'm game. Thanks again, Esmail MonkeeSage wrote: > Esmail Bonakdarian wrote: >> I would like to extract the filename with the .jpg extension. >> What is the best approach for this? > > Hi there, > > You could use raw regexps and do it yourself, but you should probably > use an HTML parser to extract HTML data. ;) > > A nice HTML parser is Hpricot [1], but it requires an extension (you > cen get it very easily via gems, see the link below). It is very easy > to use, and it's fast. > > Using Hpricot, you can do something like this: > > require 'hpricot' > require 'open-uri' > soc = open('http://utopia.utexas.edu/maps/ireland.html') > doc = Hpricot(soc) > soc.close > doc.search('//a').each { |elem| > href = elem.attributes['href'] > if not href.nil? and > ['.jpg', '.jpeg'].include?(File.extname(href)) > puts href > end > } > > Note that you can also use the built-in REXML parser [2], and do > something like: > > require 'rexml/document' > require 'open-uri' > include REXML > soc = open('http://utopia.utexas.edu/maps/ireland.html') > doc = Document.new(soc) > soc.close > doc.elements.each('//a') { |elem| > href = elem.attributes['href'] > if not href.nil? and > ['.jpg', '.jpeg'].include?(File.extname(href)) > puts href > end > } > > [1] http://code.whytheluckystiff.net/hpricot/ > [2] http://www.germane-software.com/software/rexml/docs/tutorial.html > > Regards, > Jordan >