From: Robert Klemme Date: 2004-05-12T16:48:53+09:00 Subject: Re: Strange regexp behaviour in gsub "Kristof Bastiaensen" schrieb im Newsbeitrag news:pan.2004.05.12.01.24.42.706400@vleeuwen.org... > On Wed, 12 May 2004 02:12:19 +0200, Florian Gross wrote: > > > Kristof Bastiaensen wrote: > > > >> Yes, that would clarify the situation, but is it the correct > >> behaviour? I would think that (?!a)a doesn't mean the same > >> character, but consecutive ones. Because it doesn't consume > >> the character, it effectively is the character 'before' the > >> match (if any). The other behaviour wouldn't make sense, > >> because (?!a)b is then exactly the same as b. > > > > I think that it's the intended behavior. Just use /(?!a).b/ if you want > > to consume the character. > > Hi, > You are right, I looked it up in the manual, and there it was. The > term zero-width-look-ahead pretty much says it all. I must have > gotten the definition all wrong. > > > > > Thinking about this, it is indeed possible to implement fixed-width > > look-behind -- interesting. > > I was thinking more about something like variable-width look-between :-) > Meaning for example a(?^\w+)b would match any a(.*)b if (.*) is > not equal to (\w+) IMHO that's not generally possible with regular expressions. You'll always have to define positively things that should match. Exclusion character classes are just a means of convenience but this does not extend to complete (sub) expressions. For example: to match a.*a where the part in the middle does not contain only b's (i.e. matches /b+/) you can do: /a(.*[^b].*)?a/ irb(main):004:0> rx=/a(.*[^b].*)?a/ => /a(.*[^b].*)?a/ irb(main):005:0> rx === "aa" => true irb(main):006:0> rx === "aba" => false irb(main):007:0> rx === "acba" => true Regards robert