From: Rob Doug Date: 2009-10-23T02:12:01+09:00 Subject: Re: memory leak John W Higgins wrote: > Evening Rob, > > May I ask why you need threads and that level of complication? Are you > really that sensitive towards speed that it really matters if you simply > forked processes that died after 100 downloads and just started another > worker? It would seem to be easier to use a nice simple independant > message > queue and just fire up workers that grab messages and download away. > I know it's not the sexy option but what exactly are you gaining by > using > threads here? Much easier to get a nice single threaded worker tuned up > and > tight then what you appear to be going through here. Your solution is good and simple, but there is one problem... I'll have 1-2 mil links in my database to crawl. With only one process it will take months. Beside URLs in my list sometimes are very slow, and response time may take up to 10-20-30 second. I know the best practice should would be multi-threads with asynchronous sockets, but this is too complicated for me right now, maybe in the next versions. >In fact - I might > argue > that you would be better with a downloader and separater parser > processes > that worked independantly via messages. I using mechanize, as I know it parse page right after download with Nokogiri, so separating threads would be good when I start to use net/http or sockets. Unfortunately right now I should use mechanize. -- Posted via http://www.ruby-forum.com/.