From: Ashley Moran Date: 2006-07-27T18:13:47+09:00 Subject: Re: For performance, write it in C On Thursday 27 July 2006 05:23, Francis Cianfrocca wrote: > Eliminating the dupes is the only scary thing I've seen here. What's the > absolute smallest piece of data that you need to look at in order to > distinguish a dupe? (If it's the whole line, then the answer is 16 > bytes- the length of the MD5 hash ;-)) That's the critical working set. > If you can't get the Ruby version fast enough, it's cheap and easy to > sort through 15,000,000 of them in C. Then one pass through the sorted > set finds your dupes. I've never found a consistently-fastest performer > among Ruby's several different ways of storing sorted sets. > > Make sure that your inner loop doesn't allocate any new variables, > especially arrays- declare them outside your inner loop and re-use them > with Array#clear. Nice MD5 trick! I'll remember that. Fortunately the files that need duplicate elimination are really small, so I won't need to resort to that. But I'll remember it for future reference. > > Obviously I haven't seen your code or your data, but if the Ruby app is > memory-bus-bound, then this approach may make your problem worse, not > better. Hadn't thought of that, good point... > Good luck. I recently got a Ruby program that aggregates several LDAP > directory-pulls with about a million entries down from a few hours to a > few seconds, without having to drop into C. It can be done, and it's > kindof fun too. Next time I get a morning free I might apply some of the tweaks that have been suggested. Might be interested to see how much I can improve the performance. Cheers Ashley -- "If you do it the stupid way, you will have to do it again" - Gregory Chudnovsky