From: Daniel Harple Date: 2006-04-27T02:02:59+09:00 Subject: Re: Problem with URI.parse On Apr 26, 2006, at 5:09 PM, Jeremiah Dodds wrote: > Alright, I think I may have stumbled upon a bug, correct me if I'm > wrong. > > I just wrote a script to pull linked-to files of a certain type of of > webpages. Granted, I'm still pretty new to Ruby. . . this really > has me > stumped though. I've googled around, and looked through all the > docs on > the relevant classess and I'm not getting anywhere with this. You really shouldn't be trying to parse html with regular expressions, there are a few libraries to do this available. Instead of using an external program (wget), you can also download the URI in ruby. Trying your code I get: (URI::InvalidURIError)/uri/common.rb:432:in `split': bad URI(is not URI?): here which seems reasonable. Here is an example using RubyfulSoup: require 'open-uri' require 'fileutils' require 'uri' require 'rubygems' # http://docs.rubygems.org/ require 'rubyful_soup' # http://www.crummy.com/software/RubyfulSoup/ -- sudo gem install rubyful_soup class Grabber def initialize(uri, file_types=[]) @uri = uri @file_types_re = %r{#{Regexp.union(*file_types)}$} end def grab_files find_uris.each do |link| begin data = open(link) { |a| a.read } file_path = link.host + link.path FileUtils.mkdir_p(File.dirname(file_path)) open(file_path, 'wb') { |f| f.write(data) } rescue Exception => e $stderr.puts "#{e.class}: #{e}" end end end def find_uris soup = BeautifulSoup.new(open(@uri) { |f| f.read }) soup.find_all('a') { |a| a['href'] =~ @file_types_re }.map do |a| uri = URI.parse(a['href']) # Create an absolute uri uri.host ? uri : URI.join(@uri, uri) end end end Grabber.new('http://google.com', %w{html}).grab_files __END__ -- Daniel