From: Bill Kelly Date: 2009-10-23T03:00:24+09:00 Subject: Re: memory leak From: "Rob Doug" > > Your solution is good and simple, but there is one problem... I'll have > 1-2 mil links in my database to crawl. With only one process it will > take months. > Beside URLs in my list sometimes are very slow, and response time may > take up to 10-20-30 second. > I know the best practice should would be multi-threads with asynchronous > sockets, but this is too complicated for me right now, maybe in the next > versions. Might want to look at EventMachine: http://rubyeventmachine.com/ example (from lib/em/protocols/httpclient2.rb) # === Usage # # EM.run{ # conn = EM::Protocols::HttpClient2.connect 'google.com', 80 # # req = conn.get('/') # req.callback{ |response| # p(response.status) # p(response.headers) # p(response.content) # } # } You could create multiple such concurrent connections. And if your OS supports it, EventMachine will make use of kqueue or epoll behind the scenes for efficient handling of large numbers of I/O handles. This is a simplistic example, but should work in principle: sites = %w(aaa.com bbb.com ccc.com ddd.com eee.com fff.com) # etc... conns = sites.map {|dn| EM::Protocols::HttpClient2.connect(dn, 80)} conns.each do |conn| req = conn.get('/') req.callback do |response| p(response.status) p(response.headers) p(response.content) end end This all takes place in a single thread. Regards, Bill