From: Francis Cianfrocca Date: 2006-08-18T00:16:33+09:00 Subject: Re: Ruby, memory and speed ------=_Part_97583_17110885.1155827790628 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: 7bit Content-Disposition: inline On 8/17/06, Guillaume Marcais wrote: > > I have a script that aggregates data from multiple file, store it all > in a hash, and then emit a summary on standard input. The input files > (text files) are fairly big, like 4 of about 50Mb and 4 of about 350Mb. > The hash will grow to about 500 000 keys. The memory footprint of the > ruby process as reported by top is above 2 Gigs. This is a perfect example of what I've noticed many times: Ruby's performance is perfectly fast and acceptable until the working set gets a certain (not terribly large) size, then it falls off a cliff. GC perhaps has something to do with it, but I suspect that's only a small part of the problem. Before I spend a lot of time understanding your algorithm: is there a specific reason why you need to keep the entire set in memory? You say you're generating summary output at the end of the run, but can you accumulate your summary in an object of fixed size? Or does your summary depend on some kind of transformation on the entire set (like a numeric sort)? If so, there are things you can do to improve the algorithm so you're not keeping a large working set in memory. ------=_Part_97583_17110885.1155827790628--