From: Robert Klemme Date: 2010-01-18T17:35:50+09:00 Subject: Re: Strange Encoding Behavior 2010/1/18 Lui Kore : > I know String#encode doesn't change the original string, but the result > is encoded. > > To understand the problem, you should try in a gbk/shift-jis environment > with some Chinese or Japanese path. > > The point is: > For some path p1 and p2, > when p1 == p2 and p1.encoding == p2.encoding, > p1.encode('utf-8') == p2.encode('utf-8') is not always true. > > To describe it in a "encode!" version: > For some path p1 and p2, > when p1 == p2 and p1.encoding == p2.encoding, > p1.encode!('utf-8') > p2.encode!('utf-8') > p1 == p2 is still not always true Apparently I misread your posting, sorry. Is UTF-8 capable of representing those Japanese or Chinese characters? I believe I remember Matz saying that UTF-8 is insufficient to properly represent Japanese characters. If this is the case then I guess all bets are off and you get undefined behavior. Although it might be desirable to get the same garbage it may not be worthwhile to ensure this purely for efficiency reasons. Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/