From: Dick Davies Date: 2006-07-31T19:34:58+09:00 Subject: Re: How to md5 a file? On 31/07/06, Timothy Goddard wrote: > Ben Johnson wrote: > > Basically I want to generate an md5 hash from considerably large files > > to determine if they are exactly the same. Is there a better way to do > > this besides comparing md5 hashes? > > > > Thanks for your help. > > > > -- > > Posted via http://www.ruby-forum.com/. > > I conducted a few tests to compare the performance of different > comparison methods. I tested using string comparison, the zlib > library's crc32 checksum, and the Digest::MD5 hash. The file is > iterated over in chunks and the 1K, 10K, etc refer to the size of the > chunks. There is also a whole file measure for each of them. > > The test files were identical Ogg Vorbis audio files just below 8MB in > size (identical files should give worst-case performance). Times are > for 100 repetitions. > > Rehearsal ------------------------------------------------------- > ...... removed for brevity > -------------------------------------------- total: 214.900000sec > > user system total real > String 1K 13.400000 4.250000 17.650000 ( 10.612437) > String 10K 7.633333 4.716667 12.350000 ( 7.420777) > String 100K 7.616667 4.166667 11.783333 ( 7.071255) > String Whole 7.300000 6.433333 13.733333 ( 8.260925) > CRC32 1K 16.700000 4.466667 21.166667 ( 12.774677) > CRC32 10K 9.833333 4.600000 14.433333 ( 8.769574) > CRC32 100K 9.383333 4.166667 13.550000 ( 8.129907) > CRC32 Whole 9.016667 6.333333 15.350000 ( 9.221654) > MD5 1K 26.833333 4.833333 31.666667 ( 19.087961) > MD5 10K 16.133333 4.333333 20.466667 ( 12.327322) > MD5 100K 15.216667 4.083333 19.300000 ( 11.703880) > MD5 Whole 14.633333 6.333333 20.966667 ( 12.634441) > > Notice that using MD5 is significantly slower than normal string > comparison. This also demonstrates that there are few performance gains > between 10KB buffers and 100KB buffers, indicating that somewhere in > the 10K range would be a good buffer size for the memory/performance > tradeoff. > > Of course if you really need speed you may want to code in C and > improve these times further, but a comparison rate of almost 100MB per > second isn't too shabby. > > Here's the test code for those interested: > > require 'zlib' > require 'digest/md5' > require 'benchmark' > > def step_blocks(file_a, file_b, block_size) > until file_a.eof? Won't this return true for cases where the files are of different sizes but not necessarily identical (e.ge file_b = file_a with trailing stuff)? > a = file_a.read(block_size) > b = file_b.read(block_size) > yield a, b > end > end -- Rasputin :: Jack of All Trades - Master of Nuns http://number9.hellooperator.net/