From: Dmitrii Dimandt Date: 2006-06-17T18:17:05+09:00 Subject: Re: Unicode roadmap? On 6/16/06, Juergen Strobel wrote: > On Fri, Jun 16, 2006 at 03:39:00AM +0900, Austin Ziegler wrote: > > On 6/15/06, Juergen Strobel wrote: > > [ snip essentially accurate information ] > > > > >UTF-8 encodes every Unicode code point as a variable length sequence > > >of 1 to 4 (I think) bytes. > > > > It could be up to six bytes at one point. However, I think that there > > is still support for surrogate characters meaning that a single glyph > > *might* take as many as eight bytes to represent in the 1-4 byte > > representation. Even with that, though, those are rare and usually > > user-defined (private) ranges IIRC. This also doesn't deal with > > (de)composed glyphs/combining glyphs. > > No. According to wikipedia, it is upt to 4 bytes for plain UTF8 for > all characters. Only Java may need more than that because of their use > of UTF16 surrogates and special \0 handling in an intermediary step. See > > http://en.wikipedia.org/wiki/UTF-8 > > > > > >Currently Unicode requires 21 bit, but this has changed in the past. > > > > Yes. Unicode went from 16-bit (I think) to 32-bit to 21-bit. Well, there is the official http://unicode.org/ site no one has mentioned so far. There's all sorts of technical information on Unicode. http://www.unicode.org/reports/index.html Including the latest version: http://www.unicode.org/versions/Unicode4.1.0/