From: Do One Date: 2009-02-22T12:50:48+09:00 Subject: Re: encoding problem with tr() and hash keys (1.9.1) Problem described is under modern ruby 1.9.1 in utf-8 environment. 7stud -- wrote: >> ruby 1.8.2 > > Whoops. Make that: ruby 1.8.6 (2007-03-13 patchlevel 0) [i686-linux] irb(main):001:0> h = {"a" => 1, "\u0101" => 2} => {"a"=>1, "u0101"=>2} See? It even dont understand unicode escape sequence \uXXXX. Do One wrote: > Please help to understand solution to this problem (ruby 1.9.1): > > In utf-8 environment I do: > > irb(main):121:0> h = {"a" => 1, "\u0101" => 2} > => {"a"=>1, "ā"=>2} > irb(main):122:0> h.key? "a".tr("z", "\u0101") > => false <--- wrong! > irb(main):123:0> h.key? "\u0101".tr("z", "\u0101") > => true > > So after I change utf-8 string without extended chars in it with tr(), > where second character set is having extended chars, new string is not > found in hash. > > Boths string are same in Marshal encoding: > > irb(main):124:0> Marshal.dump "a".tr("\u0101", "\u0101") > => "\x04\bI\"\x06a\x06:\rencoding\"\nUTF-8" > irb(main):126:0> Marshal.dump "a" > => "\x04\bI\"\x06a\x06:\rencoding\"\nUTF-8" > > > Question is how I should code using tr() that new string will be found > in hash? > > And I think this is bug in ruby, because it is completely not expected > behavior. -- Posted via http://www.ruby-forum.com/.