From: Julian 'Julik' Tarkhanov Date: 2006-08-01T00:24:43+09:00 Subject: Re: state of unicode support On 28-jul-2006, at 21:13, why the lucky stiff wrote: > Ruby itself also understands UTF-8 regular expressions to a > degree. Using the > 'u' modifier. Many Ruby-based UTF-8 hacks are based on the idea of: > str.scan(/./u), which returns an array of strings, each string > containing a > multibyte character. (Also: str.unpack('U*').) Which is actually useless because this breaks your string between codepoints, not between characters. ICU4R currently resolves this, as well as a library posted on ruby-talk a while ago (with proper text boudnary handling). -- Julian 'Julik' Tarkhanov please send all personal mail to me at julik.nl