From: TAKAHASHI Masayoshi Date: 2001-12-03T23:43:33+09:00 Subject: [ruby-talk:27350] Re: Unicode string for the standard library ? Rik Hemsley wrote: > > the utf-8 representation works fine for me, as regexes do support them. > > Is UTF-8 really ok ? Ruby's string is byte sequence, but Ruby's regex supports UTF-8 string. For example, /./u matches one UTF-8 character(one code unit). ###------------------------------------------------------------ # Ruby UTF-8 handling sample ## Notice: In this sample, 'character' means 'code unit'. ## split by character str = "soci\303\251t\303\251" # societe; 'e' is with acute p str.split(//u) #=> ["s", "o", "c", "i", "\303\251", "t", "\303\251"] ## character match str.scan(/.(.)./u){|c| p c[0] } #=> "o" ## 2nd character #=> "\303\251" ## 5th character ## length p str.split(//u).length # => 7 ## character length p str.split(//n).length # => 9 ## byte length p str.length # => 9 ## byte length ## replace str2 = str.gsub(/\xC3\xA9/u,"e") p str2 #=> "societe" ## other replace str3 = str2.gsub(/e/u,"\xC3\xA9") p str3 #=> "soci\303\251t\303\251" ###------------------------------------------------------------ Regards, TAKAHASHI 'Maki' Masayoshi E-mail: maki@open-news.com