From: Venherm.Borchers@... (Venherm Borchers) Date: 2002-03-20T05:26:27+09:00 Subject: Re: Why is Ruby so slow? - Solution Dear colleagues, thanks for your attention and comments. The following suggestions were made, but were obviously not true: (1) It is not a "slow file reading" problem, as you can see from the timings: 2 secs for loading the data into memory, then 48 secs for processing. (2) It strongly appears not to be a Ruby 1.7 versus 1.6 problem as almost the same timing behavior can be seen in 1.6.5 and 1.7.2 implementations. (3) It may be a Windows problem, but some tests I did on a (indeed much slower) Linux machine indicate similar differences in running times. The solution ------------ Nobu Nakada pointed out that the following line might be causing the problem, and it indeed does: nitms = (@col[j]-['']).nitems Here I relied on the 'set difference' that returns an array of only unique elements which then are counted. I have to warn you: this kind of set operation is _unbelievably_ slow! It need not be, as the set difference can be implemented as a simple hash, counting the number of keys in the end. Doing this myself I got the following runnung times for Ruby 1.7.2 as compared to the Python program, applying it to 50,000 lines of data and 1.6 Mill. data elements: ________Total___________loading...______analyzing...____ Ruby 86 sec 17 sec 69 sec Python 48 sec 7 sec 41 sec Now these are comparable execution times I can live with, keeping in mind to do some optimization later and going ahead to implement some data analysis tools in Ruby. Thanks. Hans Werner. P.S.: By the way, counting 'nitms' up in each block made it slightly slower for a few seconds, and the same with this line: nitms = @col[j].nitems - @col[j].grep(/^$/).nitems besides not giving exactly the right result.