From: sa 125 Date: 2008-11-27T23:39:19+09:00 Subject: Re: Performance issues with large files -- ruby vs. python : I can't really put the code here since it's on the company's intranet. I use the fastercsv and mysql libraries. Basically I want to grab the first and last record of every date/company/product combo, and store it's row info. The core processing is done through if statements @main_hash = {} csv = FasterCSV.open(file_path, "r", :headers => true) #...code below is in loop: for row in csv... if not @main_hash.keys.member?(date) @main_hash[date] = {} @main_hash[date][company] = {} @main_hash[date][company][prod] = {} @main_hash[date][company][prod] = row_values else if not @main_hash[date].keys.member?(company) @main_hash[date][company] = {} @main_hash[date][company][prod] = {} @main_hash[date][company][prod] = row_values else if not @main_hash[date][company].keys.member?(prod) @main_hash[date][company][prod] = {} @main_hash[date][company][prod] = row_values end end end # loop ends This is basically the part of the code that runs slow. I keep track of progress in percentage (file position / file size) throughout the loop. I should mention I extract the row values into loop variables, like date/company/prod using the csv headers: date = row['Date'], etc. the row_values variable is an array containing all the relevant parameters from the row. The @main_hash variable obviously takes up some memory. There are a couple of if-statements, but not much else. That's pretty much all I can think about. Thanks! -- Posted via http://www.ruby-forum.com/.