From: Robert Klemme Date: 2004-08-19T17:15:54+09:00 Subject: Re: Regexp scanning with MatchData (Re: multiple regexp matches) "Austin Ziegler" schrieb im Newsbeitrag news:9e7db911040818075512cd5a01@mail.gmail.com... > On Wed, 18 Aug 2004 19:31:01 +0900, Robert Klemme wrote: > > "Austin Ziegler" schrieb im Newsbeitrag > > news:9e7db91104081713254f2eb39e@mail.gmail.com... > >> str = ' ... ' > >> re = /(<(\/?)span> )/i > >> > >> str.scan(re) > >> # => [[" ", ""], [" ", "/"], [" ", "/"]] > >> > >> matches = [] > >> str.scan(re) do > >> matches << Regexp.last_match > >> end > >> > >> matches.each do |match| > >> match.captures.each_with_index do |capture, ii| > >> soff, eoff = match.offset(ii + 1) > >> puts %Q("#{capture}" #{soff} .. #{eoff}) > >> end > >> end > > While that works, isn't it ridiculous that one has to resort to a > > class method ("Regexp.last_match")? I mean, there should rather be > > something like > > > > /o/.each( "foo" ) do |md| > > # md is MatchData > > end > > There's a simple solution, and I'll probably open an RCR about this > if others agree with it. String#scan, #sub, and #gsub should yield > MatchData objects, not Strings. There are probably others, but those > are the ones that come to mind. This *will* break some code, > unfortunately, but that can be mitigated by adding #to_str. IMO, > this will make #gsub much easier to deal with, as you won't have to > resort to either Regexp.last_match or $[0-9] variables to be able to > work with captures. My Regexp.last_match call only presumes that > Regexp.last_match is actually threadsafe, whereas we know that the > ugly Perlish $ variables are threadsafe. I think this is an > acceptable level of incompatibility because of the use of #to_str > and the amount of flexibility that would be gained. As far as I > know, it wouldn't require *that* big a change, because for > Regexp.last_match to work, there must still be a MatchData object > *somewhere*. > > What do you think? I like the functionality very much, but I'd prefer to *not* change the behavior of String#scan, #sub, and #gsub. I'd rather have Regexp#scan(str, &block), Regexp#sub(str, replace=nil, &block) and Regexp#gsub(str, replace=nil, &block) that yield MatchData if there is a block. There might be other names but since the behavior is quite similar to those methods in String these names are propably good. The only drawback I can see is that they might cause confusion ("Which were the ones that yielded MatchData?"), but IMHO people can cope with this - especially since old behavior does not change. (Personally I would find it easy to remember that Regexp <-> MatchData and String <-> String or Array of String.) Kind regards robert