From: brabuhr@... Date: 2009-10-23T03:05:44+09:00 Subject: Re: memory leak On Thu, Oct 22, 2009 at 1:12 PM, Rob Doug wrote: > John W Higgins wrote: > > Evening Rob, > > > > May I ask why you need threads and that level of complication? Are you > > really that sensitive towards speed that it really matters if you simply > > forked processes that died after 100 downloads and just started another > > worker? It would seem to be easier to use a nice simple independant > > message > > queue and just fire up workers that grab messages and download away. > > I know it's not the sexy option but what exactly are you gaining by > > using > > threads here? Much easier to get a nice single threaded worker tuned up > > and > > tight then what you appear to be going through here. > Your solution is good and simple, but there is one problem... I'll have > 1-2 mil links in my database to crawl. With only one process it will > take months. > Beside URLs in my list sometimes are very slow, and response time may > take up to 10-20-30 second. > I know the best practice should would be multi-threads with asynchronous > sockets, but this is too complicated for me right now, maybe in the next > versions. I'm sure he didn't mean (only) 1 process but several processes each running a subset of the input space. It could also have the benefit of utilizing multiple processing cores on one machine or across multiple machines: http://raa.ruby-lang.org/project/rq/ http://blade.nagaokaut.ac.jp/cgi-bin/scat.rb/ruby/ruby-talk/298739 http://pragmatig.wordpress.com/2009/08/13/ruby-parallel-processing-made-simple/