From: Garance A Drosihn Date: 2002-01-24T14:32:42+09:00 Subject: Re: Minor thoughts on 'Find' module At 3:45 PM +0900 1/23/02, Yukihiro Matsumoto wrote: >Hi, Greetings. >In message "Minor thoughts on 'Find' module" > on 02/01/23, Garance A Drosihn writes: > >|I've seen some messages about ways to get ruby to do the depth-first >|trick, but I do not remember those as being options to 'Find', which >|seems to me to be the logical place to have the option. > >Could you propose the possible API of the "new Find"? Let me start out with a few "design goals": 1) there is work that Find must perform in order to do what it does (such as doing a 'stat' on every filename to find out if that file is a directory). Let us take the maximum advantage of whatever work Find has to be doing anyway. 2) One of the parameters to Find is a code-block, which (of course) can do any kind of arbitrarily complicated amount of work. We should not add parameters to Find for it to do "extra work" (any work besides the work it *has* to do), when it is so much more flexible for that "extra work" to be done in the code-block. 3) There is some work which Find does not have to do, but which is much much easier for Find to do than for the code-block to do. The main thing that comes to mind here is altering the order in which filenames are processed. 4) Something I am not sure of... does the current Find have any problems with recursive Find calls? For instance, is there any interaction between the two Finds in: Find.find(".") do |outerName| statinfo = File.lstat(aname) if statinfo.directory? Find.find("/other/dir") do |innerName| if xyzzy(outerName, innerName) # Does this Prune effect the outer Find? # if not, how *can* the inner code-block # trigger a prune of the outer code-block? Find.prune end end end end So, I think I want something like: Find.findx(sort=File::NO_SORT, parent_first=true, breadth_first=false, *srcdirs ) { |aFileName, aFindStat| block } I used "Find.findx" in case this new API is too radical a change from the current Find.find. "findx" is meant to indicate "find - extended". Other values for sort= would be File::DIR_FIRST File::DIR_LAST and perhaps File::ALPHABETIC Probably the most noteworthy change is adding the 'aFindStat' parameter to the code-block. This would allow the code block to find out (and set?) various values that are known by the Find processing. For instance: aFindStat.depth --> current "depth" of the search from the "srcdir" which is currently being processed. This would let the code-block implement "max_depth" or "min_depth" processing if it wanted to. I think it's reasonable for Find to keep track of this value, because I expect it must be easier for Find to keep track of the level than it would be for the code-block to determine it based on the value of aFileName. (particularly when Find is given multiple source-dirs, such as "." and "/mnt/home/gad"). aFindStat.fileStat --> in many of my Find's (in either ruby or perl), the first thing I do is a stat() call on the value of aFileName. This just duplicates the stat() call that File procesing itself had to do, so why not give the code-block the file-status info that Find already has? [ Note: this would be FileStat from an lstat() call, so the code-block can tell if aFileName is a symlink] aFindStat.sourceDir --> the source directory which Find is currently processing. This might be useful when Find was called with multiple srcdirs. aFindStat.parentDir --> the parent directory for aFileName. I am not sure how often this would be useful, but I assume that Find processing does already have that value so it might as well be available to the code-block. I assume this value would be relative to the sourceDir which Find is currently processing, and not necessarily an absolute directory name. aFindStat.basename --> This is another value that the code-block could figure out easily enough, but I suggest the option only because I assume that the Find processing probably has the value anyway. These two values may save the code-block from having to care about what the separator character is on different OS platforms ("/" vs "\" vs ":"). And two values which could be set by the code-block: aFindStat.prune --> alternate to Find.prune, but it's clearer which "find" would be altered. This is always initialized to "false" before calling the code-block, and the code-block could change it to "true". aFindStat.followlink -> if the current aFileName is a symlink, and if that symlink also points to a directory, then Find processing should treat the symlink as if *it* were a directory, and descend into the directory pointed to by the symlink. This is always initialized to "false" before calling the code-block, and the code-block could change it to "true" to modify the behavior of find. There are a few details of how this should work that I have not decided on yet... In particular, should Find repeat the call to the code-block, but this time send in aFileStat which matches the directory being pointed to? I think we would want something like that, particularly when thinking how this would interact with the sort= option on the call to Find. This would allow the code block to implement "follow_symlinks", but also with some finer-grain control. Such as "follow any symlinks which point outside of the current sourceDir". I believe that with this API, the combination of Find.xfind and the user's code-block will be able to implement anything and everything that are implemented by the standard Unix 'find' commands, and without too much clutter in the calling sequence. I do not know if this is consistent with other standard ruby classes, but I think it's good enough to provide ideas which are useful (although perhaps in some other form). -- Garance Alistair Drosehn = gad@eclipse.acs.rpi.edu Senior Systems Programmer or gad@freebsd.org Rensselaer Polytechnic Institute or drosih@rpi.edu