From: Kevin Howe Date: 2004-08-18T06:46:00+09:00 Subject: Re: multiple regexp matches Awesome that works great thank you. I have to wonder why Ruby doesn't have this built in, it's simple enough to add a method that returns a list of MatchData objects as follows: class MultiRegexp < Regexp def matches(str) str.scan(self) do yield Regexp.last_match end end end str = ' ... ' re = MultiRegexp.new('(<(\/?)span>)', true) re.matches(str) { |i| capture = i.captures[0] start,stop = i.offset(0) puts "\"#{capture}\" starts at #{start}, ends at #{stop}" } An even nicer alternative would be to add a Regexp::MULTIMATCH constant: str = ' ... ' re = Regexp.new('(<(\/?)span>)', Regexp::MULTIMATCH) matches = re.match(str) Just a thought :) "Zach Dennis" wrote in message news:41226FE4.4060108@mktec.com... > Ah...thanks for the clarification David. I was mistaken. > > Sorry for the confusion Kevin. > > Zach > > David A. Black wrote: > > >Hi -- > > > >On Wed, 18 Aug 2004, Zach Dennis wrote: > > > > > > > >>According to rdoc you are mistaken. > >> > >>I also think you are mistaken: > >> > >>#!/usr/bin/ruby > >>t = "This is my 1 text" > >> > >>re = /([^\s]*\s).*(\d)(\s.)/ > >>md = re.match( t ); > >>puts md.offset(0); > >>puts "" > >>puts md.offset( 1 ); > >>puts "" > >>puts md.offset( 2 ); > >>puts "" > >>puts md.offset( 3 ); > >> > >> > >>It returns the correct offsets of the matches. offset(0) being the whole > >>regex, offset(1) does the first subexpression, offset(2) does the second > >>subexpression. It works. > >> > >> > > > >The problem is that Kevin wanted to scan a string more than once with > >the same regex: > > > > str = "abc abc abc" > > re = /(\w+)/ # not /(\w+) (\w+) (\w+)/ > > > >re will scan against str three times. The difficulty is getting hold > >of the offsets of all the matches from all three times, in relation to > >the total length of the string. > > > >Someone will probably post a simple or elegant solution; in the > >meantime, here's mine: > > > > def find_offsets(str,re) > > offsets = [] > > first = 0 > > of = [0,0] > > > > loop do > > break unless m = re.match(str[first..-1]) > > break if m.captures.empty? > > m.captures.each_with_index do |c,i| > > of = m.offset(i+1) > > res = [c, [of[0]+first, of[1]+first ]] > > yield res if block_given? > > offsets << res > > end > > first += of[0] > > end > > > > offsets > > end > > > > # Little test: > > > > str = ' ... ' > > re = /(<(\/?)span>)/i > > > > puts str > > (str.size/9).times { print "0123456789" } > > puts; puts > > > > find_offsets(str,re).each do |capture, (start, stop)| > > puts "\"#{capture}\" starts at #{start}, ends at #{stop}" > > end > > > > # Output: > > ... > > 0123456789012345678901234567890123456789 > > > > "" starts at 14, ends at 20 > > "" starts at 15, ends at 15 > > "" starts at 23, ends at 30 > > "/" starts at 24, ends at 25 > > "" starts at 31, ends at 38 > > "/" starts at 32, ends at 33 > > > > > >David > > > > > > > > >