From: Vincent Isambart Date: 2007-06-17T19:22:37+09:00 Subject: Re: Ruby Changes Its Mind About Non-Word Characters > I agree that the � and � are logically quotes (after all I'm French > and they are used a lot in French), but what I said is that the > Unicode support for Ruby 1.8 is poor (in fact except split(//) you > can't do much) so I am not that suprised by the fact it does not work. I just checked in ruby 1.8.6's code what was done and it is indeed quite simple. Everything is based on whether a character "is a letter" or not, and "is a letter" is defined by (I simplified it a bit): - if the characters is an ASCII character (ASCII code <= 127) or the multibyte character mode is ASCII (the default if you did not change it), the result is what says the standard C function isalnum (is alpha numeric) - if the characters code is bigger that 127 and the multibyte character mode is not ASCII, it "is a letter" if the character takes more than one byte. And in UTF-8, _all_ characters with a code bigger than 127 take at least 2 bytes. So in ruby 1.8 in UTF-8 mode, all non ASCII characters are considered a letter. Yes it is not a good thing but think that this is not so simple, knowing if a Unicode character is a letter or not requires a better regexp engine, including tables with the attributes of all Unicode characters. But for that you will need to use ruby 1.9. Tim Bray had reasons to say that Unicode support in ruby 1.8 sucked. And Matz knows it, that's what ruby 1.9 will be better on this point.