From: nobu.nokada@... Date: 2002-01-09T03:31:18+09:00 Subject: Re: Meaning of \G (was: Re: Efficient "lexing" in Ruby (maybe an RCR?)) At Tue, 8 Jan 2002 02:01:13 +0900, Matt Armstrong wrote: > >> I liked your idea at first, but this method is actually slower in my > >> case. Why? String#index with an offset does not support \A or ^ in > >> the regexp. > > > > Use \G instead. \A always matches the beginning of the string. > > Yes, this works, thanks. Strangely, the "String#index with \G" > version is still a little bit slower than a version that keeps > chopping up the string with $POSTMATCH and using Regexp#~= with \A. > I'll have to investigate. Hmmm, you might test with -K option? $ cat /tmp/linecount_postmatch.rb src = IO.read(ARGV.shift) cnt = 0 while /\A.*\n/ =~ src cnt += 1 src = $' end p cnt $ cat /tmp/linecount_index.rb src = IO.read(ARGV.shift) cnt = 0 pos = 0 while pos = src.index(/\G.*\n/, pos) cnt += 1 pos += $&.length end p cnt $ time ruby /tmp/linecount_postmatch.rb eval.c 9170 real 0m11.477s user 0m8.210s sys 0m3.280s $ time ruby /tmp/linecount_index.rb eval.c 9170 real 0m0.200s user 0m0.200s sys 0m0.010s $ time ruby -Ke /tmp/linecount_postmatch.rb eval.c 9170 real 0m11.680s user 0m8.660s sys 0m3.010s $ time ruby -Ke /tmp/linecount_index.rb eval.c 9170 real 0m17.721s user 0m17.570s sys 0m0.040s I guess regex.c:re_adjust_startpos() is a bottleneck. Since String#scan doesn't adjust start position on the way, it's fast as String#index without -K option. It needs backward scan on multibyte string. > But the meaning of \G is not clear. Matz' book says "match point > where last match finished" -- but for String#index that definition is > confusing. Does \G for String#index just mean "anchor at beginning of > offset?" It can be said, "match point where last match finished, or start position for first match". -- Nobu Nakada