From: Curt Sampson Date: 2002-08-01T19:53:07+09:00 Subject: Re: Unicode in Ruby now? On Thu, 1 Aug 2002, Tobias Peters wrote: > ... > We also need rules how to combine strings with different encoding then. > concatenating two strings encoded in koi8-r and iso-8859-1, respectively, > may only be possible when the result is encoded in some unicode > representation. Yeah. This is getting into complex nightmare city. That's why I'd prefer to have the basic system just work completely in Unicode. One could have a separate character system (character and string classes, byte-stream to char converters, etc.) to work with this tagged format if one wished. > Are there any other character sets of relevance that are not part of > unicode yet? Yeah. There are tons of obscure Japanese characters that are not in and will never be in Unicode, some of which exist in various other character sets. In particular there's Mojikyo (http://www.mojikyo.org/) which is at 80,000 characters and growing. > I understand that Mojikyo will not be folded into unicode due to > political reasons. Not just political reasons, but practical reasons. Unicode is designed to work if you restrict yourself to using only 16-bit chars, and I expect most programs are going to limit themselves to that. So even if it were folded in to the extension space, most people wouldn't use it. cjs -- Curt Sampson +81 90 7737 2974 http://www.netbsd.org Don't you know, in this new Dark Age, we're all light. --XTC