From: Ryan Davis Date: 2011-03-23T06:46:21+09:00 Subject: Re: How could I make the Ruby 1.9 string ignore the invalid utf-8 byte sequence in split? This sounds like it might be a legitimate bug. Can you file a ticket on redmine with this code sample? On Mar 22, 2011, at 09:09 , Joey Zhou wrote: > I am working with Chinese character radicals. > I came across a radical which has a codepoint "\uE839". > > ----- > > # encoding: utf-8 > [ STDIN, STDOUT, STDERR ].each do |stdio| > stdio.set_encoding( 'gbk', 'utf-8' ) > end > char = "\uE839" > puts char # Encoding::UndefinedConversionError > > ----- > f.rb:7:in `write': U+E839 from UTF-8 to GBK > (Encoding::UndefinedConversionError) > from f.rb:7:in `puts' > from f.rb:7:in `puts' > from f.rb:7:in `
' > > But Perl works > > ----- > use utf8; > use open ":encoding(gbk)", ":std"; > > $char = "\N{U+E839}"; > print $char; > ----- > It prints out what I want--a chinese character radical > > -- > Posted via http://www.ruby-forum.com/. >