From: Austin Ziegler Date: 2006-06-18T11:34:51+09:00 Subject: Re: Unicode roadmap? On 6/17/06, Julian 'Julik' Tarkhanov wrote: > On 18-jun-2006, at 0:21, Austin Ziegler wrote: >>> I am unaware of unsolveable problems with Unicode and Eastern >>> languages, I asked specifically about it. If you think Unicode is >>> unfixably flawed in this respect, I guess we all should write off >>> Unicode now rather than later? Can you detail why Unicode is >>> unacceptable as a single world wide unifying character set? >>> Especially, are there character sets which cannot be converted to >>> Unicode and back, which is the main requirement to have Unicode >>> Strings in a non-Unicode environment? >> Legacy data and performance. > Yes, you will spend those cycles to count the letters in my language > RIGHT :-)) (evil grin) It's actually the most common case when apps > damage strings in my language - their authors wanted to be smart and > _conserve_. And yes, normalization etc. is complex and you DO need to > have a case-conversion table in memory. Please do have one (Ruby > doesn't). I think you're overthinking the problem. Let's consider the guarantees that an m17n String would make: * #size and #length would return the number of glyphs * #[] would return glyphs Presumably, in Regexen with an m17n String, \w would indicate only "word" glyphs. Other guarantees *would* be made along that line. Therefore, if your input data is UTF-8, anything that deals with #size, #length, and character-based indexing *will just work*. The same will apply to SJIS or any other encoding. The number of times that people are dealing with mixed-encoding data is vanishingly small, and even when a developer must, they will probably use a Unicode encoding to deal with that. But if you're using SJIS, you're just going to want use *that*. That's what the m17n String is about. It's not about dictating a single encoding, but enabling people to use Strings intelligently. > No offense, just observation. I agree -- we *need* full Unicode support. But not at the cost of legacy code pages in favour of Unicode. It's not always appropriate. -austin -- Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ * austin@halostatue.ca * http://www.halostatue.ca/feed/ * austin@zieglers.ca