From: David Vallner Date: 2006-01-11T15:11:47+09:00 Subject: Re: Ruby, Unicode - ever? Austin Ziegler wrote: >>(current state of UTF8 in Ruby, even with regexps, is too far away >>from proper Unicode support, don't try to cheat me, that it's OK and >>enough, it is not!) >> >> > >For 99% of cases, in fact, is *is* sufficient. What do you think is >missing? > > > How would the regexp engine match multibyte UTF8 characters that have what is ASCII whitespace as one of the lower bytes? Or how would /\w{2,4}/ react to a single three-byte UTF-8 character? I didn't yet stumble upon this in the rather spartan kcode documentation, does the UTF8 support for Japanese input cater for these perks? >>Well, I think, String class must be REMOVED from Rite. Instead, two >>incompatible classes must be introduced: ByteArray and Text with >>well-separated semantics and behaviour. Else it will never end but >>eventually crash into crap ruins someday... >> >> > >You're welcome to submit an RCR on it. I am 99.999% certain it will be >shot down, though. > > > Shot, hung, drawn, and quartered probably. The slight abmuguity of strings might be baffling for people with a Java or similar background, but Dmitry, if your arguments are to hold water, I want you to give examples or real-life code where it isn't possible to tell when a String is used to store binary data, and when it is storing text, in a situation where this distinction is necessary to process the string. Otherwise, carry on rambling emptily into a Notepad window. As for the legacy encoding support, I -wish- I saw that more often. Try getting any work done on an English Windows XP, with Slovak regional settings for the odd ancient non-Unicode tool, and a German keyboard, and you start wanting to access the encoding of the consoles real soon. David Vallner