From: ahoward Date: 2003-04-25T02:31:23+09:00 Subject: [OT] Re: grep and regular expressions in ruby On Fri, 25 Apr 2003, Warren Brown wrote: > Ara, > > > notsmall = '^[^s]*$|^[^s]*(?:s(?!mall)[^s]*)+$' > > This caught my eye as being a little redundant. To illustrate, if we > let X = '[^s]*' and Y = '(?:s(?!mall)[^s]*)', this can be rewritten as: > > notsmall = "^#{X}$|^#{X}#{Y}+$" > > Which makes it a little easier to see that this is the same as: > > notsmall = "^#{X}#{Y}*$" > > Undoing the substitutions: > > notsmall = '^[^s]*(?:s(?!mall)[^s]*)*$' > > However, making a distinction between the 's' and the 'mall' still > bothered me. At first I tried simply: > > notsmall = '^(?!small)*$' > > Unfortunately, this only matches a zero length string since the (?!re) > syntax does not consume any characters. So, I added a '.' to consume the > current character: > > notsmall = '^(?:(?!small).)*$' > > I believe this behaves the same as your original regular expression, but > (IMHO) is much clearer. However, I have no idea what the performance > implications are. > > I hope someone finds this interesting... it it interesting - i've filed it in my 'interesting' folder ;-) however, it doesn't work for the examples my program was attempting to solve? if you notice the program has the pattern .... :cat_and_notsmall => %r/(?:(?:#{cat}.*#{notsmall})|(?:#{notsmall}.*#{cat}))+/, .... for some reason, your pattern doesn't seem to work if nested that way. in fact, now that i look at it i'm not sure *how* mine does either. try plugging yours in the program and see what i mean - ideas on this? in any case - i think it IS clear that matching 'not word' is not only possible, it might even be easier than i made it out to be ;-) -a -- ==================================== | Ara Howard | NOAA Forecast Systems Laboratory | Information and Technology Services | Data Systems Group | R/FST 325 Broadway | Boulder, CO 80305-3328 | Email: ara.t.howard@fsl.noaa.gov | Phone: 303-497-7238 | Fax: 303-497-7259 ====================================