From: Julian 'Julik' Tarkhanov Date: 2006-08-01T00:35:13+09:00 Subject: Re: state of unicode support (OT) On 31-jul-2006, at 17:10, Tim Bray wrote: > Unicode characters have named properties. "L" means it's a > letter. There are sub-properties like Lu and Ll for upper and > lower case. There are lots more properties for things like being > numbers, being white-space, combining forms and particular > properties of Asian characters and so on. Tremendously useful in > regexes, particularly for those of us round-eye gringos who are > prone to write [a-zA-Z] and think we're matching letters, which > we're not. If you don't support properties, you don't support > Unicode. That's one of the reasons why you _need_ tables when working with Unicode, and you _will_ spend memory on them. What Ruby does now is nowhere near, and Matz wrote that he didn't unclude complete tables for Oniguruma in 1.9 yet. With proper regex support other funky things become posslbe, for instance {all_cyrillic_letters} in a regex etc. -- Julian 'Julik' Tarkhanov please send all personal mail to me at julik.nl