From: Brian Candler Date: 2012-01-14T17:52:08+09:00 Subject: Re: Ruby1.9: Encoding problems (how to use #force_encoding ?) Pedro G. wrote in post #1040715: > "That's correct. Transcoding tries to *transcode* (replace characters > one > at a time), and these high characters in ASCII-8BIT have no Unicode > equivalents." > > It isn't true Yes, it *is* true, because it's exactly what the Ruby encoding "ASCII_8BIT" means. It allows you to use \x80 to \xFF without defining what character set those are in. Hence these characters cannot be transcoded, since it's undefined what they are. (Also, why are you resurrecting a 2-year-old thread?) > , in fact, \xc characters are Unicode code points (in UTF-8 > encoding) and not ASCII-2 characters. You have a Unicode String with > UTF-8 encoding and encoding incorrectly set to ASCII-2, to solve this > problem try this: What do you mean by ASCII-2? Standard ASCII is only a 7-bit character set. There are a whole bunch of 8-bit extensions to ASCII, e.g. ISO-8859-1, Windows-1252 etc. They all define different character sets for \x80 to \xff. The encoding "ASCII_8BIT" makes no assertion about what these high characters are. > begin > str.encode! Encoding::UTF_8 if str.encoding != Encoding::UTF_8 > rescue Encoding::UndefinedConversionError > #string incorrectly encoded try force > str.force_encoding Encoding::UTF_8 > end That's wrong, and just shows you don't understand the problem. -- Posted via http://www.ruby-forum.com/.