From: Zach Dennis Date: 2004-08-18T05:52:34+09:00 Subject: Re: multiple regexp matches Ah...thanks for the clarification David. I was mistaken. Sorry for the confusion Kevin. Zach David A. Black wrote: >Hi -- > >On Wed, 18 Aug 2004, Zach Dennis wrote: > > > >>According to rdoc you are mistaken. >> >>I also think you are mistaken: >> >>#!/usr/bin/ruby >>t = "This is my 1 text" >> >>re = /([^\s]*\s).*(\d)(\s.)/ >>md = re.match( t ); >>puts md.offset(0); >>puts "" >>puts md.offset( 1 ); >>puts "" >>puts md.offset( 2 ); >>puts "" >>puts md.offset( 3 ); >> >> >>It returns the correct offsets of the matches. offset(0) being the whole >>regex, offset(1) does the first subexpression, offset(2) does the second >>subexpression. It works. >> >> > >The problem is that Kevin wanted to scan a string more than once with >the same regex: > > str = "abc abc abc" > re = /(\w+)/ # not /(\w+) (\w+) (\w+)/ > >re will scan against str three times. The difficulty is getting hold >of the offsets of all the matches from all three times, in relation to >the total length of the string. > >Someone will probably post a simple or elegant solution; in the >meantime, here's mine: > > def find_offsets(str,re) > offsets = [] > first = 0 > of = [0,0] > > loop do > break unless m = re.match(str[first..-1]) > break if m.captures.empty? > m.captures.each_with_index do |c,i| > of = m.offset(i+1) > res = [c, [of[0]+first, of[1]+first ]] > yield res if block_given? > offsets << res > end > first += of[0] > end > > offsets > end > > # Little test: > > str = ' ... ' > re = /(<(\/?)span>)/i > > puts str > (str.size/9).times { print "0123456789" } > puts; puts > > find_offsets(str,re).each do |capture, (start, stop)| > puts "\"#{capture}\" starts at #{start}, ends at #{stop}" > end > > # Output: > ... > 0123456789012345678901234567890123456789 > > "" starts at 14, ends at 20 > "" starts at 15, ends at 15 > "" starts at 23, ends at 30 > "/" starts at 24, ends at 25 > "" starts at 31, ends at 38 > "/" starts at 32, ends at 33 > > >David > > >