From: Austin Ziegler Date: 2006-03-14T15:07:01+09:00 Subject: Re: unicode in ruby On 3/13/06, Anthony DeRobertis wrote:> >> UTF-8 can take multiple octets to represent a character. So can> >> UTF-16, UTF-32, and every other variation of Unicode.> > This last statement is true only because you use the term "octet."> You're correct; that isn't what I meant to say. Something along the> lines of the following is better worded:>> UTF-8 can take more than one octet to represent a> character; UTF-16 can take more than two; UTF-32> more than four; etc. No. UTF-32 does not have surrogates. Unicode is perfectlyrepresentable in either 20 or 21 bits. A single character is *always*representable in a uint32_t sized space with UTF-32. POSIX is outdated and needs to be scrapped or fixed. Preferably theformer. Preferably by people who know what they're doing -- and notthe folks behind the GNU libc. -austin--Austin Ziegler * halostatue@gmail.com * Alternate: austin@halostatue.ca