From: Robert Klemme Date: 2007-12-30T21:15:01+09:00 Subject: Re: using reg expr with array.index On 28.12.2007 02:17, Esmail wrote: > MonkeeSage wrote: >> >> def no_dups(path) >> IO.read(path).split(" >").uniq.join(" >") >> end >> fixed = no_dups("testfile") >> puts fixed > > One more quick questions (ha .. see, that's what you get > for being so helpful) - please feel free to ignore this. > > For the above solution which I really like, is there an > easy way to get the duplicate records? (I'd like to display > the name lines ie the ones that start with > as a possible > check of what I am eliminating from the original data). > > I know how to do this if I reread the file again and traverse > it but that's certainly not an efficient way to do this. > > I am doing a lot of reading on Ruby right now, so I'll > may come across the solution, so only reply if you are bored :-) > > > > (ps: I suppose if there was a to_set and to_array functionality > in Ruby - for all I know there is - it would have yet provided > another approach to solve the original problem) Esmail (btw, is that a real name?), since you are dealing with files I would like to revisit this. Since files a) are slower to read and b) are potentially large - especially larger than main memory - you might want to look for different solutions. Basically, unless you need the file's contents otherwise you should try to avoid having to store the whole file in memory at one point in time and strive to process the whole file as seldom as possible (ideally only once). Here are two typical approaches: 1. if you know the file is ordered File.open("foo") do |io| last = nil io.each do |line| if line == last $stderr.puts "Duplicate line no #{io.lineno}" else puts line last = line end end end 2. if the file can be unordered File.open("foo") do |io| dups = Hash.new 0 io.each do |line| line.freeze # optimization for Hash key c = (dups[line] += 1) if c > 1 $stderr.puts "#{c}. occurrence of line at #{io.lineno}" else puts line end end end Advantage of both these approaches is that the file has to be read only once. However, the second solution still has the whole file's contents in memory at some point in time. If you know more about your data (for example, that repetitions always occur within n lines) you can create more efficient algorithms (with the mentioned restriction it is sufficient to just remember the last n lines, similar to the first approach). An alternative would be to store a more compact representation of lines, e.g. as a MD5 hash and do the lookups of the second solution based on hash codes of lines. However this approach is less strict, i.e. although unlikely there might be lines reported as duplicates because they accidentally yield the same hash code despite having different content. But if you need to manually edit the file anyway this approach might be sufficient for large files. Once you dive into the matter, all sorts of interesting problems surface. :-) But it is generally good to know the nature of the data; with this often great optimizations can be done. Kind regards robert