From: Brian Candler Date: 2008-11-28T01:45:13+09:00 Subject: Re: Performance issues with large files -- ruby vs. python : sa 125 wrote: > if not @main_hash.keys.member?(date) @main_hash.keys will create an array of all the keys, which will be expensive when it's large, and member? will do a linear search, which is also expensive. You can replace with: if not @main_hash.has_key?(date) or more simply if not @main_hash[date] > @main_hash[date] = {} > @main_hash[date][company] = {} > @main_hash[date][company][prod] = {} > @main_hash[date][company][prod] = row_values The third line does nothing, because it's replaced by the fourth line. I think all you need is: @main_hash[date] = { company => { prod => row_values } } > else > if not @main_hash[date].keys.member?(company) > @main_hash[date][company] = {} > @main_hash[date][company][prod] = {} > @main_hash[date][company][prod] = row_values > else > if not @main_hash[date][company].keys.member?(prod) > @main_hash[date][company][prod] = {} > @main_hash[date][company][prod] = row_values > end > end > end You can rewrite this as above too. But looking at this, I think you can replace *all* this code with just the following three lines: @main_hash[date] ||= {} @main_hash[date][company] ||= {} @main_hash[date][company][prod] ||= row_values Note that a ||= b is the same as a = a || b, which will assign b to a only if a is nil or false. > This is basically the part of the code that runs slow. I keep track of > progress in percentage (file position / file size) throughout the loop. There is also the ruby profiler, which you can turn on/off where needed, or just run the whole lot with ruby -rprofile (beware: makes your code run *much* slower) Regards, Brian. -- Posted via http://www.ruby-forum.com/.