From: matz@... (Yukihiro Matsumoto) Date: 2002-03-25T15:07:16+09:00 Subject: Re: Unicode in Regexp followup Hi, In message "Re: Unicode in Regexp followup" on 02/03/25, Sean Russell writes: |By the way, after I did some research, I found that all characters above |0x7F are remapped; therefore, I'd expect problems with string matching UTF8 |regexps with the entire character range of 0x80-0xFF... but Ruby only gives |trouble with the range of 0xF0-0xFD. I find this rather odd. By the definition of UTF-8, 0x80-0xEF at the first byte of multibyte sequence are invalid, so Ruby treats them as if they are single byte characters. matz.