From: Rob Biedenharn Date: 2007-07-25T02:27:07+09:00 Subject: Re: Regex Question On Jul 24, 2007, at 12:50 PM, seijin@gmail.com wrote: > I recently started studying Ruby (1.8.6) and things are going well. > > However, something bothered me when I got to the regular expression > section. This is my first time deal with regular expressions directly > so I'm completely new to things. I understand the idea behind regular > expressions and most of the usage but two of them stuck out for me. > > I'm studying the 2nd edition Pragmatic book and using their > show_regexp method... > > def show_regexp(a, re) > if a =~ re > "#{$`}<<#{$&}>>#{$'}" > else > "no match" > end > end > > ...hopefully I don't get in trouble posting that method here. So here > are my two problems... > > First, why does ... show_regexp("banana", /(an)*/) ... not match > "anan" ? I thought it was a greedy algorithm that tried to match as > much as possible? Well, it would but the * is 'zero or more times' and it can match zero "an"s before it ever needs to get past the 'b'. If you change the '*' to a '+' and force 'one or more times', then it has to find the first "an" and will then greedily consume the next "an" also. > Second, how does ... show_regexp("Mississippi", /(\w+)\1/) ... work? > Why in the world does it match "ississ" rather than returning no > match? I think most of my problem with this one is not understanding > the underlying logic used when doing pattern matching with back > references. Does it go through and check "Mississippi", "Mississipp" > down to "M" and then "i", "is", "iss", etc? Like trying all possible > combinations in a lock? I would be extremely grateful if someone > would do a short step-by-step of how it matches "ississ". I > understand what the "\1" does, I just don't understand how the first > part even gets to the first "iss". > > Thanks a lot! The '+' is greedy so I think it tries \w+ as Mississippi and finds that there's no following copy, backs up to try Mississipp and finds that 'i' isn't a copy, and so on back to 'M' before moving past 'M' to 'ississippi', 'ississipp', 'ississip', etc. until 'iss' has a following copy and a match is declared. Think of it this way: Until the regexp gets to the end, it doesn't know that the string isn't "MississippiMississipp". -Rob Rob Biedenharn http://agileconsultingllc.com Rob@AgileConsultingLLC.com