From: Raul Parolari Date: 2007-12-02T18:15:25+09:00 Subject: Re: Negate a character sequence in a regular expression? unknown wrote: > I also tweaked Tanaka's solution a little bit after a trip back to the > pickaxe book. I added the (?: to keep the group from generating > backreferences, Great > and I added an alternate 'dog' to the negative assertion. Just using a non-greedy expression removes this complication (as you realized later), although it made the following even more interesting. > Here is my attempt at breaking it down, for anyone who is interested: > /cat(?:(?!cat|dog).)*dog/ > 1 2 3 4 5 6 > > 1. The engine starts scanning through the string until matches 'cat'. > 2. At the first position... This analysis was great (I had missed Tanaka's solution). > ...In any event, I think I still like Daniel's solution better, > because I can look at it and feel fairly certain that it will do > exactly what it should do. Josh I think both solutions are indestructible. I was curious to benchmark them; here, they showed interesting differences: a) for one-line sentences, where each substring cat...dog contains a few words, Tanaka's solution is approx 25% faster. b) for long paragraphs where cat...dog are separated by (say) 2 dozen words, Daniels' solution becomes faster by 20-25% (for longer intervals, the advantage can go up to 50% and more). c) For long paragraphs, where cat and dog were separated by only a few words, the performance is almost identical. So, it is not so much the length of the string, but the frequency of appearance of the keywords that is important. [Note: you need to add the /m modifier if your strings are paragraphs. In fact, it may be better to always write /m, so that you do not need to touch the regexp]. The reason for the performance discrepancies is clear: for short gaps between the keywords (ie cat, dog), the trio scan+map+gsub exacts a toll. But for long paragraphs with many words between the keywords, it is instead the negative lookahead on 'cat' (at every character) which suffers. You may want to take this in account, depending on the type of text configurations you expect. In any case, both methods seem technically perfect. All the best, and thanks for that great summary, Raul -- Posted via http://www.ruby-forum.com/.