From: Robert Klemme Date: 2009-04-29T20:02:40+09:00 Subject: Re: Random Access using IO#pos in code blocks 2009/4/29 Brian Candler : > Unfortunately, I don't think string#scan will give you offsets into the > strings found. > > In ruby 1.8 you can write this: > >  pos = 0 >  while md = /\b\w+\b/.match(para[pos..-1]) >    word = md[0] >    puts "Match #{word} at #{para_pos+pos+md.begin(0)}" >    pos += md.end(0) >    ... >  end > > In ruby 1.9 (but not 1.8.6/1.8.7), Regexp.match takes a start pos, so > you could optimise it to this: > >  pos = 0 >  while md = /\b\w+\b/.match(para, pos) >    word = md[0] >    puts "Match #{word} at #{para_pos+md.begin(0)}" >    pos = md.end(0) >    ... >  end String#scan is likely faster than manually matching portions with #match. In both versions of Ruby you can do this to get the /character/ offset: irb(main):001:0> s=%{foo bar baz} => "foo bar baz" irb(main):002:0> s.scan(/\w+/) { p $`.length } 0 4 8 => "foo bar baz" > However in ruby 1.9 the offsets used will be in terms of number of > characters, not number of bytes. It would be up to you to convert this > back into byte offsets into the file, if that's what you're after. This is an important point to remember! Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/