From: Randy Kramer Date: 2005-03-21T23:20:35+09:00 Subject: Re: Iterating through a string and removing leading characters On Monday 21 March 2005 04:44 am, Robert Klemme wrote: > Some remarks: > > - The comparison between 5 and 6 does not seem fair, as you iterate in 6 > but not in 5. Oops, for a minute I thought I had really screwed up (like by not doing 5 the 10000 times). AFAIK, I don't need to iterate (through the string) in 5 as the RE is not anchored to the beginning of the string--it still checks the entire string (line) for that pattern, which is what I try to achieve in 6 by the iteration through the string and check a character then invoke RE approach. I am sure that my Ruby code to do that is not the best, and I may learn something by making it better, but I agree with your conclusion / recommendation (below) at least for the time being (although I do plan to play with str::scan and StringScanner at least a little bit (I presume they do similar things, but perhaps the 2nd is optimized somehow, particularly if I "require C" or whatever)). > - You don't use String#scan or #split which you are likely to need in > practice, because you want to sift through complete documents and want to > treat all occurrences. I need to let that sink in a bit. In general I do want to treat all occurrences, but I plan to scan a line (actually a paragraph) at a time, and some things can only occur at the beginning of a line, so those would only be checked at the beginning of a line. > - The differences between the check-first-char approach and the pure RE > approach are so insignificant that I'd not bother using the more complex > code. I'd stick with pure RE based approaches and only try to optimize if > performance is bad. (You mentioned premature optimization already... :-)) Thanks! I pretty much agree at this time, although I want to play a little bit with StringScanner. regards, Randy Kramer