From: Daniel DeLorme Date: 2007-12-06T11:31:07+09:00 Subject: Re: Unicode in Regex MonkeeSage wrote: > Here is a micro-benchmark on three common string operations (split, > index, length), using bytestrings and unicode regexp, verses native > utf-8 strings in 1.9.0 (release). That's nice, but split and index do not operate using integer indexing into the string, so they are rather irrelevant to the topic at hand. They produce the same results in ruby1.8, i.e. uni_split==reg_split and uni_index==reg_index. I also stated that the point of regex manipulation is to *obviate* the need for methods like index and length. So a more accurate benchmark might be something like: reg_chars N/A N/A N/A ( N/A ) uni_chars 0.130000 0.000000 0.130000 ( 0.193307) ;-) > Ps. BTW, in case there is any confusion, bytestrings aren't going > away; you can, as you see above, specify a magic encoding comment to > ensure that you have bytestrings by default. Yes, it's still possible to access bytes but it's not possible to run a utf8 regex on a bytestring if it contains extended characters: $ ruby1.9 -ve '"abc" =~ /b/u' ruby 1.9.0 (2007-12-03 patchlevel 0) [i686-linux] $ ruby1.9 -ve '"日本語" =~ /本/u' ruby 1.9.0 (2007-12-03 patchlevel 0) [i686-linux] -e:1:in `
': character encodings differ (ArgumentError) And that kinda kills my whole approach. Daniel