From: Eivind Eklund Date: 2005-01-14T00:15:39+09:00 Subject: Re: The face of Unicode support in the future On Fri, Jan 14, 2005 at 12:03:36AM +0900, Yukihiro Matsumoto wrote: > Hi, > > In message "Re: The face of Unicode support in the future" > on Thu, 13 Jan 2005 23:42:55 +0900, Eivind Eklund writes: > > |> Yes, and the comparison is always false unless > |> > |> * encodings of the two strings are both ASCII compatible > |> * they have same (7 bits) ASCII character sequence > | > |This feels wrong to me. My hunch tells me it often will work > |right in test situations, and then fall down in edge cases > |during production. I'd be much happier with something that > |just failed directly (ie, the strings would always be different) > |than something that often-but-not-always returns the right > |value. > > It was so in my first prototype, but the compatibility between plain > ASCII and UTF-8 (and other Japanese encodings) is used so often, so > widely, above condition was preferred. Can you imagine any edge case? Yes. I'll use one that has bit me in other directions: Input comes from a database and a web form. The subsystems provide strings with different encoding. During tests, programmers use english for testing (most programmers I know prefer using english for all code development including tests, even if their native language is norwegian/german/etc). The tests work correctly. When the system is set in production, things start to fall apart. I've been bitten by the encoding issue between Integer and String due to the combination of a database and a text file. This was slightly icky. I find it likely that having the comparison work 99% of the time and fail the rest would cause much more pain. I'd almost wish comparison of strings with different encoding to raise an exception, but it would be in conflict with the way other comparisons work, and I'm not sure which behaviour (exception vs always returning false) would be most useful for strings anyway. Eivind.