From: Clifford Heath Date: 2002-08-05T12:58:27+09:00 Subject: Re: Unicode in Ruby now? Curt Sampson wrote: > These statements are both very wrong. I was deliberately being "reinterpretive", but what I said is the effective truth. If you want to do Unicode-3 correctly and simply, then a 32 bit internal representation is the right one - surrogates simply recreate exactly the same problems of variable-length encoding that plagued earlier efforts at providing a simple way (for the programmer) to code correctly. Externally, a more compact encoding is needed (utf-8 or utf-16 are valid choices), but internally, UTF-16 is bogus in the extreme. Even internally, *if appropriately hidden* behind an API that only exposes whole characters, a more compact encoding (such as I've recently described) can be worthwhile. You seem to be so wedded to the Java/Unicode model that you can't see out of the hole into which you've dug yourself. -- Clifford Heath