From: Kirk Haines Date: 2010-08-07T02:20:24+09:00 Subject: Re: comparing two arrays, too slow On Fri, Aug 6, 2010 at 9:51 AM, joe chesak wrote: > I don't have access to the database, but everyday I get a csv dump of all > live customer listings. > There are around 15,000 live listings. > A record contains 3 fields: id, name, description > Everyday some new accounts are created while others are deleted. > Everyday I want to compare yesterday's dump with today's dump and print > the full record of each terminated record and each new record. > I am achieving this with the script below. > That last line makes the script run very slow. > > Is there a more elegant way to compare two > arrays|hashes|sets|FasterCSVtables while carrying a few fields along for the > ride? > > > WORKING CODE: > require 'rubygems' > require 'fastercsv' > > sundays_dump = FasterCSV.read("./sunday.csv") > mondays_dump = FasterCSV.read("./monday.csv") > > sundays_ids = sundays_dump.collect {|row| row[1]} > mondays_ids = mondays_dump.collect {|row| row[1]} > > newaccount_ids = mondays_ids - sundays_ids > terminated_ids = sundays_ids - mondays_ids > > sundays_dump.each {|row| puts 'delete,'+row[0]+','+row[1] if > terminated_ids.include? row[1]} > mondays_dump.each {|row| puts 'create,'+row[0]+','+row[1] if > newaccount_ids.include? row[1]} The answer to, "it's too slow", is often, "don't do that, then." If comparing arrays is too slow, don't compare arrays. Instead, do something that's fast. Looking up data in a Hash is fast. And since your data appears to be keyed by unique ids (though your code is written with variable names assuming it is looking at ids, but with array lookups that look at the names, which is confusing and might be a bug), it is easy to stuff that data into a hash for later lookup. ----- require 'rubygems' require 'fastercsv' def extract_data(dumpfile) dump = FasterCSV.read(dumpfile) ids = [] data = {} dump.each do |row| ids << row[0] data[row[0]] = row end [ids,data] end sundays_dump = FasterCSV.read("./sunday.csv") mondays_dump = FasterCSV.read("./monday.csv") sunday_ids, sunday_data = extract_data("./sunday.csv") monday_ids, monday_data = extract_data("./monday.csv") newaccount_ids = monday_ids - sunday_ids terminated_ids = sunday_ids - monday_ids terminated_ids.each do |id| row = sunday_data[id] puts "delete,#{id},#{row[1]}" end newaccount_ids.each do |id| row = monday_data[id] puts "create,#{id},#{row[1]}" end ----- If this were my script, I'd further clean it up so that all of your logic is in a simple library. Then either create a generic executable that takes the file names as arguments, or write your scripts with hard coded file names, but that just calls into your library with those names. It'd be much cleaner unless you are absolutely sure that this code will never be more than a one-off hack. Kirk Haines