From: Ashley Moran Date: 2006-07-27T06:16:25+09:00 Subject: Re: For performance, write it in C On Jul 26, 2006, at 9:31 pm, Francis Cianfrocca wrote: > Ashley was talking about 1GB+ datasets, iirc. I'd love to see an > in-memory data structure (Ruby or otherwise) that can slug a few of > those around without breathing hard. And on most machines, you're > going > through the disk anyway with a dataset that large, as it thrashes your > virtual-memory. So why not take advantage of the tunings that are > built > into the I/O channel? > > If I'm using C, I always handle datasets that big with the kernel vm > functions- generally faster than the I/O functions. I don't know > how to > do that portably in Ruby (yet). I think the total data size is about 1.5GB, but the individual files are smaller, the largest being a few hundred GB. The most rows in a file is ~15,000,000 I think. The server I run it on has 2GB RAM (an Athlon 3500+ running FreeBSD/amd64, so the hardware is not really an issue)... it can get all the way through without swapping (just!) The processing is pretty trivial, and mainly involves incrementing some ID columns so we can merge datasets together, adding a text column to the start of every row, and eliminating a few duplicates. The output file is gzipped (sending the output of CSV::Writer through GzipWriter). I could probably rewrite it so that most files are output a line at a time, and call out to the command line gzip. Only the small files *need* to be stored in RAM for duplicate removal, others are guaranteed unique. At the time I didn't think using RAM would give such a huge performance hit (lesson learnt). I might also look into Kirk's suggestion of FasterCSV. If all this doesn't improve things, there's always the option of going dual-core and forking to do independent files. However... the script can be run at night so even in its current state it's acceptable. It will only need serious work if we start adding many more datasets into the routine (we're using two out of a conceivable 4 or 5, I think). In that case we could justify buying a faster CPU if it got out of hand, rather than rewrite it in C. But that's more a reflection of hardware prices than my wages :) I have yet to write anything in Ruby was less than twice as fast to code as it would have been in bourne-sh/Java/whatever, never mind twice as fun or maintainable. I recently rewrote an 830 line Java/ Hibernate web service client as 67 lines of Ruby, in about an hour. With that kind of productivity, performance can go to hell! Ashley