From: TAKAHASHI Masayoshi Date: 2001-12-12T02:26:34+09:00 Subject: [ruby-talk:28251] Re: Unicode string for the standard library ? Hi, Sorry to response too late. martin@v.loewis.de (Martin v. Loewis) wrote: > > Ruby's string is byte sequence, but Ruby's regex supports > > UTF-8 string. For example, /./u matches one UTF-8 > > character(one code unit). > > How can I do Unicode ranges? This won't work > > str = "soci\303\251t\303\251" > p str.gsub(/[\303\251-\303\254]/u,"?") > > I expect to get one replacement character per Unicode character, but > actually get two of them. Maybe you're right. Unfortunately, there are two problems in current implementation of Regexp in Ruby. * There's no support of \x notation for multibyte character in range. So, you should write: str = "soci\303\251t\303\251" p str.gsub(/[#{"\303\251"}-#{"\303\254"}]/u,"?") or str = "soci\303\251t\303\251" p str.gsub(/#{"[\303\251-\303\254]"}/u,"?") * Moreover, there's a bug handling 0x80-0xff characters in UTF-8. so you should get new Ruby sources from CVS and use it, or apply a patch (from [ruby-dev:15360] by K.Kosako-san): --- regex.c.orig Tue Oct 30 17:42:21 2001 +++ regex.c Wed Dec 5 15:07:01 2001 @@ -477,7 +477,7 @@ } while(0) #define WC2MBC1ST(c) \ - ((c<0x100)?(c):((current_mbctype != MBCTYPE_UTF8)?(((c)>>8)&0xff):utf8_firstbyte(c))) + ((current_mbctype != MBCTYPE_UTF8) ? ((c<0x100) ? (c) : (((c)>>8)&0xff)) : utf8_firstbyte(c)) static unsigned int utf8_firstbyte(c) Regards, TAKAHASHI 'Maki' Masayoshi E-mail: maki@open-news.com