From: Julian 'Julik' Tarkhanov Date: 2006-06-19T00:29:46+09:00 Subject: Re: Unicode roadmap? On 18-jun-2006, at 13:08, Michal Suchanek wrote: > > But quite a few people here look like they do know. I do not know much > about regexes but I can imagine just about any other string operation. > And the current regexes already do operate on multiple encodings. Oh, lord... Have you at least tried that to make such assumtpions? In other words, tell me, can Ruby's regexes cope with the following: /[а-я]/ /[а-я]/i or something like this: http://rubyforge.org/cgi-bin/viewvc.cgi/icu4r/samples/demo_regexp.rb? revision=1.2&root=icu4r&view=markup > > > And how that leads to the conclusion that there should be only one > encoding? Very simply - I use many pieces of software written in many languages all the time, with non-Latin text. I know that when they want to get "historically compatible" problems arise. And the software that settles on Unicode internally or somehow enforces it on the programmer usually works best (all Cocoa and all C#. And to a certain extens yes, Java). > >> >> Bluntly put, I am selfish and I don't believe in the "saving grace" >> of the M17N (because I just can't wrap it around my head and I sure >> as hell know it's going to be VERY complex). > > That's the point. If it is wrapped into the string class you do not > have to wrap it around your head. This is rather naive. > > And that is eaxctly why a fixed encoding is bad. If strings can be > encoded in any way there is no point i religious discussions which > encoding you like the most. Yes, it just becomes hard and error prone to process them. > > It is JustGoodEnouhg for most cases but not for all. It is not useless > for CJK, just suboptimal because of the Han unification. And it also > does not try to include the historic characters. I think this thread is going to end the same as the one in 2002 did. -- Julian 'Julik' Tarkhanov please send all personal mail to me at julik.nl