From: Panagiotis Atmatzidis Date: 2012-11-15T00:53:07+09:00 Subject: Re: Efficient way for comparing records between 2 large files (16 million records) Hello, On 12 Νοε 2012, at 23:21 , Ruby Student wrote: > Team, > > I have two large files, about 16 million records each. > The files are sorted. > The first 13 characters are used as a key. > We get an updated file every week. > We also keep the previous week file. > > I need to compare the keys from the new file to the keys on the file from last week. If the rest of the records are the same, then I do nothing. If the keys matches but the rest of the records are different, I then have an update and I will output that record to a new file. > I was wondering if there is an efficient way to do this in Ruby. Either any built-in method or an efficient algorithm which I can implement. Not sure if you dealing with hashes or arrays but take a look at the '-' method for arrays, it's blazing fast for comparing even large arrays. > > Thank you > > > -- > Ruby Student Panagiotis (atmosx) Atmatzidis email: atma@convalesco.org URL: http://www.convalesco.org GnuPG ID: 0xE736C6A0 gpg --keyserver x-hkp://pgp.mit.edu --recv-keys 0xE736C6A0 -- The wise man said: "Never argue with an idiot. They bring you down to their level and beat you with experience."