From: Robert Klemme Date: 2005-08-04T02:16:09+09:00 Subject: Re: How do I reduce the memory usage of a script? Scott Ellsworth wrote: > In article <3jmmgnFqqsq6U1@individual.net>, > "Robert Klemme" wrote: > >> Scott Ellsworth wrote: >>> In article <8HKdneMVh9MZDUjfSa8jmA@karoo.co.uk>, >>> "daz" wrote: >>> >>>> if (!found) then >>>> puts "Unknown #{f}" >>>> end >>>> data = nil >>>> GC.start # garbage collect >>> >>> This did seem to drop the memory usage on my MacOS X 10.4.2 system. >>> >>> I will investigate the Find.find command next to see if I can get >>> rid of some recursion. An array of 5500 paths should not be _that_ >>> big, at least in comparison with four or five levels of directory >>> depth. >> >> The problem might be that the data is still around while you enter >> the recursion. If you want to verify that this is the case you can >> simply do data = nil after processing. But: You definitely need to >> throw out the recursion from propath() - otherwise you'll be >> processing directories over and over again (I smell something like >> O(n*n) here)! > > I have removed the recursion - see below. > > A question, though, Is the String.scan method I used the best way to > do the scan this block of data? Every file is going to contain the > string 'QuickTime' somewhere in the first few MB, and I want from the > last nonprintable character before it to the next nonprintable > character after. I only need to read from disk until I find that > string, and once I find it, I need only the bytes before, plus a > version number afterwards. I certainly do not need to manipulate > more than a few hundred characters around that magic string, and once > I have read, I do not need to go back. > > NB - nonprintable here is defined as [[:alnum:]_., ] The problem with your script is that it does not find "QuickTime" if your chunk reading cuts it in half (or "Q" and "uickTime" - whatever). It might be easier to just slurp in the complete file (depending on size - a few MB are no problem) and then do the scan on the single string. Also, I don't understand why you don't put QuickTime into your search RE. Kind regards robert > > work@boggle:Desktop$ time ./detectEncoding.rb > songs.txt > > real 3m30.563s > user 0m26.229s > sys 0m23.746s > > New code: > > #!/usr/bin/env ruby > require 'find' > re = /[[:alnum:]_., ]{9,}/ > Find.find("/Users/work/Music/iTunes/iTunes Music/") do |f| > if File.file?(f) && File.fnmatch("*.m4a",f) then > found = false > data = IO.read(f, 65536*8) > data.scan(re) do |string| > if (string =~ /QuickTime/) then > filename = File.basename(f) > dirname = File.dirname(f) > puts "#{string} #{dirname}" > # puts "#{string} #{dirname} #{filename}" > found = true > break > end > end > if (!found) then > puts "Unknown #{f}" > end > data = nil > GC.start # garbage collect > end > end > > Scott