From: Curt Sampson Date: 2002-08-04T17:47:22+09:00 Subject: Re: Deprecation and Unicode On Sun, 4 Aug 2002, Ned Konz wrote: > Unicode is *defined* to be representable using 2-byte numbers. There > are provisions for using sequences of two (or more now?) of these > numbers to represent characters outside the basic 64K or so. Right. > There > are few applications that deal fully with these extended characters > ("surrogate pairs"?); Java (for instance) doesn't. "Java" deals with it just fine; if you do what the Unicode spec. tells you to do. You decide what level of support you need and write it in. Many common activities need no extra support for dealing with surrogate pairs. > In a program, representing Unicode by 2-byte numbers is probably ideal > unless you're very space-constrained (in which case UTF-8 is best) UTF-8 is best if you're using ISO-8859-1, yes. However, in Japanese for example, UTF-8 will take up *more* space than UTF-16. cjs -- Curt Sampson +81 90 7737 2974 http://www.netbsd.org Don't you know, in this new Dark Age, we're all light. --XTC