From: Joey Zhou Date: 2011-03-23T01:09:07+09:00 Subject: Re: How could I make the Ruby 1.9 string ignore the invalid utf-8 byte sequence in split? I am working with Chinese character radicals. I came across a radical which has a codepoint "\uE839". ----- # encoding: utf-8 [ STDIN, STDOUT, STDERR ].each do |stdio| stdio.set_encoding( 'gbk', 'utf-8' ) end char = "\uE839" puts char # Encoding::UndefinedConversionError ----- f.rb:7:in `write': U+E839 from UTF-8 to GBK (Encoding::UndefinedConversionError) from f.rb:7:in `puts' from f.rb:7:in `puts' from f.rb:7:in `
' But Perl works ----- use utf8; use open ":encoding(gbk)", ":std"; $char = "\N{U+E839}"; print $char; ----- It prints out what I want--a chinese character radical -- Posted via http://www.ruby-forum.com/.