From: Albert Schlef Date: 2011-03-23T20:59:37+09:00 Subject: Re: A question about Ruby 1.9's "external encoding" Robert K. wrote in post #988839: > What *I* find slightly puzzling is this: > > irb(main):001:0> s1 = "a" > => "a" > irb(main):002:0> s1.encoding > => # > irb(main):003:0> s2 = s1.encode 'ISO-8859-1' > => "a" > irb(main):004:0> s2.encoding > => # > irb(main):005:0> s1 == s2 > => true > irb(main):006:0> s1.eql? s2 > => true > irb(main):007:0> [s1.hash, s2.hash] > => [1003075638, 1003075638] > irb(main):008:0> [s1.hash, s2.hash].uniq > => [1003075638] > irb(main):009:0> s1.encoding == s2.encoding > => false > > Apparently only the byte representation is used for equivalence checks > and the encoding is ignored. I don't think this is true: irb(main):043:0> utf = "\u05D0" # Alef => "א" irb(main):044:0> latin = utf.dup; latin.force_encoding 'ISO-8859-1' => "�\x90" irb(main):045:0> [utf.bytes.to_a, latin.bytes.to_a] # They have the same bytes => [[215, 144], [215, 144]] irb(main):048:0> [utf.valid_encoding?, latin.valid_encoding?] # And are ok => [true, true] irb(main):046:0> utf == latin # But they aren't equal => false In your case it's good the strings are considered equal: we want to know if the letters are all the same. "a" is "a"... no matter what encoding. -- Posted via http://www.ruby-forum.com/.