From: Joseph McDonald Date: 2000-12-16T10:54:10+09:00 Subject: [ruby-talk:7371] Re: Ruby performance question Very cool David! I found that a combination of your technique with Dave's technique to be the fastest: def dt_readfile h = Hash.new file = File.new("keys") while line = file.gets record = line.chomp!.tr!('=', ',').split(',') h[record.shift] = record.shift until record.empty? # do stuff here... h.clear end end results: user system total real david: 8.750000 0.148438 8.898438 ( 9.019113) joe: 8.773438 0.109375 8.882812 ( 8.944328) dt: 6.960938 0.046875 7.007812 ( 7.130645) orig: 11.304688 0.093750 11.398438 ( 11.776709) (dt is the one above). The above is for 81,000 records of 7 records each totaling a little over 13MB on a pIII-550. All keys were unique (doesn't matter very much since we have that h.clear there...). I note that if I remove the h.clear in the method above, the process grows to 52 megs and takes this long to run: user system total real dt: 78.132812 0.195312 78.328125 ( 82.569231) (that's a total of 81_000 * 7 hash entries...). Not sure how good or bad that is... regards, -joe > -----Original Message----- > From: David Alan Black [mailto:dblack@candle.superlink.net] > Sent: Friday, December 15, 2000 5:21 PM > To: ruby-talk ML > Subject: [ruby-talk:7367] Re: Ruby performance question > > > On Sat, 16 Dec 2000, Dave Thomas wrote: > > > OK, this is tacky, and probably not worth it, but it knocks about 15% > > of the run times on my box, using an 80,000 line input file with 6 > > key/value pairs per line. The idea is simply to avoid creating the > > intermediate object that contains a single key-value which is then > > split across the '='. Instead, we assume the input is well-formed and > > split into the final strings directly. > > > > > > h = Hash.new > > > > while line = gets > > a = line.chomp!.tr!('=', ',').split(',') > > 0.step(a.size-2, 2) do |i| > > h[a[i]] = a[i+1] > > (Oh, for Array#to_h! Wouldn't: "h = line.to_h" be nice?) > > If the lines are guaranteed to be well-formed, you could also do: > > while line = gets > record = line.chomp!.split(/[,=]/) > h[record.shift] = record.shift until record.empty? > h.clear > end > > which bypasses the tr! phase. And seems to pick up a bit of time -- a > quick benchmarking of the three versions, on a file of 20000 lines > (all right, I'm impatient :-) with six key-value pairs per line, looks > like this: > > user system total real > 13.200000 0.030000 13.230000 ( 13.225895) # Eric > 11.200000 0.050000 11.250000 ( 11.242669) # Dave > 9.870000 0.040000 9.910000 ( 9.918286) # David > > > I wanted to use the reverse/pop (instead of shift) technique that > Brian Feldman had suggested for interleaving, but I couldn't figure > out how to get at the key before the value. > > > David > > -- > David Alan Black > home: dblack@candle.superlink.net > work: blackdav@shu.edu > Web: http://pirate.shu.edu/~blackdav > >