From: "Matthew K. Williams" Date: 2010-08-07T00:57:33+09:00 Subject: Re: comparing two arrays, too slow On Sat, 7 Aug 2010, joe chesak wrote: > I don't have access to the database, but everyday I get a csv dump of all > live customer listings. > There are around 15,000 live listings. > A record contains 3 fields: id, name, description > Everyday some new accounts are created while others are deleted. > Everyday I want to compare yesterday's dump with today's dump and print > the full record of each terminated record and each new record. > I am achieving this with the script below. > That last line makes the script run very slow. > > Is there a more elegant way to compare two > arrays|hashes|sets|FasterCSVtables while carrying a few fields along for the > ride? > > > WORKING CODE: > require 'rubygems' > require 'fastercsv' > > sundays_dump = FasterCSV.read("./sunday.csv") > mondays_dump = FasterCSV.read("./monday.csv") > > sundays_ids = sundays_dump.collect {|row| row[1]} > mondays_ids = mondays_dump.collect {|row| row[1]} > > newaccount_ids = mondays_ids - sundays_ids > terminated_ids = sundays_ids - mondays_ids > > sundays_dump.each {|row| puts 'delete,'+row[0]+','+row[1] if > terminated_ids.include? row[1]} > mondays_dump.each {|row| puts 'create,'+row[0]+','+row[1] if > newaccount_ids.include? row[1]} Rather than go through both sets of data, I'd do something like this: terminated_ids.each do |row| puts "delete,#{sundays_dump[row][0]},#{sundays_dump[row][1]}" end Same thing for the new account (substituting monday, of course) This way you're only outputting the changed records. Matt