From: ara.t.howard@... Date: 2006-07-27T06:24:07+09:00 Subject: Re: For performance, write it in C On Thu, 27 Jul 2006, Ashley Moran wrote: > I think the total data size is about 1.5GB, but the individual files are > smaller, the largest being a few hundred GB. The most rows in a file is > ~15,000,000 I think. The server I run it on has 2GB RAM (an Athlon 3500+ > running FreeBSD/amd64, so the hardware is not really an issue)... it can get > all the way through without swapping (just!) > > The processing is pretty trivial, and mainly involves incrementing some ID > columns so we can merge datasets together, adding a text column to the start > of every row, and eliminating a few duplicates. The output file is gzipped > (sending the output of CSV::Writer through GzipWriter). I could probably > rewrite it so that most files are output a line at a time, and call out to > the command line gzip. Only the small files *need* to be stored in RAM for > duplicate removal, others are guaranteed unique. At the time I didn't think > using RAM would give such a huge performance hit (lesson learnt). > > I might also look into Kirk's suggestion of FasterCSV. If all this doesn't > improve things, there's always the option of going dual-core and forking to > do independent files. > > However... the script can be run at night so even in its current state it's > acceptable. It will only need serious work if we start adding many more > datasets into the routine (we're using two out of a conceivable 4 or 5, I > think). In that case we could justify buying a faster CPU if it got out of > hand, rather than rewrite it in C. But that's more a reflection of hardware > prices than my wages :) > > I have yet to write anything in Ruby was less than twice as fast to code as > it would have been in bourne-sh/Java/whatever, never mind twice as fun or > maintainable. I recently rewrote an 830 line Java/Hibernate web service > client as 67 lines of Ruby, in about an hour. With that kind of > productivity, performance can go to hell! i process tons of big csv files and use this approach: - parse the first line, remember cell count - foreach line - attempt parsing using simple split, iff that fails fall back to csv.rb methods something like n_fields = nil f.each do |line| fields = lines.split %r/,/ n_fields ||= fields.size if fields.size != n_fields fields = parse_with_csv_lib line end ... end this obviously won't work with csv files that have cells spanning lines, but for simply stuff it can speed up parsing in a huge way. -a -- suffering increases your inner strength. also, the wishing for suffering makes the suffering disappear. - h.h. the 14th dali lama