From: Wes Nakamura Date: 2005-01-20T10:28:14+09:00 Subject: Re: The face of Unicode support in the future On Thu, 20 Jan 2005, Yukihiro Matsumoto wrote: | |Also, if s.explode.length == s.length, same character in different encodings: | | | |1. "\x{30b9}".explode (encoding = utf16) => [ 0x30b9 ]? | |2. "\x{b930}".explode (encoding = utf16le) => [ 0x30b9 ]? | |3. "\x{e382b9}".explode (encoding = utf8) => [ 0xe3, 0x82, 0xb9 ] or [ 0xe382b9 ]? | | I'm afraid you have misunderstood. The number within the braces | should be a code point, not a sequence of bytes. | | matz. Yes, I'm getting quite confused! :) I'm looking at: http://www.unicode.org/versions/Unicode4.0.0/ch03.pdf which is not an especially easy read. Unicode codespace: 0-10ffff (I suppose you could call this the unicode character "set") Codepoint: a value in the unicode codespace Encoded character: an association or mapping between an abstract character and a codepoint Code unit: bit combination that can represent a unit of encoded text utf-8 has 8-bit code units, and utf-16 has 16-bit code units The unicode codepoint and utf-16-encoded values are the same for I believe 0-ffff. So (these are all the same character - a katakana "su"): "\x30\xb9" - unicode codespace, utf-16 encoding, codepoint 30b9 "\xe3\x82\xb9" - unicode codespace, utf-8 encoding, codepoint 30b9 "\x25\x39" - "JIS" encoding, codepoint 2539 (in JIS X0208 codespace) "\xa5\xb9" - euc-jp encoding, codepoint 2539 (in JIS X0208 codespace) "\x83\x58" - shift-jis encoding, codepoint 2539 (in JIS X0208 codespace) Should explode give the codepoint values listed above? The bottom three examples are complicated by shift-jis and euc-jp using multiple character sets (codespaces?), not just JIS X0208. Wes