From: Tim Bray Date: 2006-06-18T13:35:17+09:00 Subject: Re: Unicode roadmap? On Jun 17, 2006, at 6:52 AM, Austin Ziegler wrote: >> The internal encoding has to be decided, probably between UTF-8, >> UTF-16, and UTF-32 by the String class implementor. > > Completely disagree. Matz has the right choice on this one. You can't > think in just terms of a pure Ruby implementation -- you *must* think > in terms of the Ruby/C interface for extensions as well. Point of information: Of all the widely-used methods of encoding international strings, UTF-8 is by far the easiest to deal with in C. > Trust me on this > one: I *have* done some low-level encoding work. Additionally, even > though I might have marked a network object as "UTF-8", I may not know > whether it's *actually* UTF-8 or not until That's an incredibly important point in a networked world. One of the reasons XML has had so much success, probably more than it deserves, is that its encoding is self-descriptive. To quote Larry Wall: "An XML document knows what encoding it's in." Since HTTP headers are (sigh) known to be wrong on occasion, this is a pretty big value-add. >> - This ties Ruby's String to Unicode. A safe choice IMHO, or would we >> really consider something else? Note that we don't commit to a >> particular encoding of Unicode strongly. > > This is a wash. I think that it's better to leave the options open. > After all, it *is* a hope of mine to have Ruby running on iSeries > (AS/400) and *that* still uses EBCDIC. EBCDIC is in fact an encoding of Unicode. Just saying that it's necessary to be clear both as to what character set is being supported, and what limitations on encoding are enforced. -Tim