From: Michal Suchanek Date: 2006-06-12T12:14:48+09:00 Subject: Re: Which encoding causes fewest problems in Ruby 1.8.2? On 6/11/06, Yukihiro Matsumoto wrote:> Hi,>> In message "Re: Which encoding causes fewest problems in Ruby 1.8.2?"> on Sun, 11 Jun 2006 10:12:45 +0900, Jim Smith writes:>> |Given the above problems (or rumors), which encoding is recommended for> |use with ruby 1.8.2?> |> |UTF-8> |ISO-8859-1> |ISO-8859-15>> String and Regexp handles all of them for most of the cases. But> upper/lower case handling for non ASCII alphabets are not supported.> Use -Ku for UTF-8 and -Kn for ISO-8859-*.> Length and indexing do not work very well with utf-8. ~ $ irb -Kuirb(main):001:0> $KCODE=> "UTF8"irb(main):002:0> a='α-ω'=> "α-ω"irb(main):003:0> r=/[β-ω]/=> /[β-ω]/irb(main):004:0> a.length=> 5irb(main):005:0> a[0..0]=> "\316"irb(main):006:0> a[0..1]=> "α" Fortunately, the regexps work. irb(main):007:0> a =~ r=> 3 So you could use a.scan /./ to calculate length or index characters in a string. irb(main):008:0> a.scan /./=> ["α", "-", "ω"]irb(main):009:0> (a.scan /./).length=> 3 Michal