From: Martin Elzen Date: 2004-05-04T23:28:44+09:00 Subject: Re: Strange behaviour of Strings in Range Hi everyone. I've looked up the references Robert Klemme found by searching for I18N on the Ruby-talk archives, I also did the same myself, and also read up some more about Unicode. One of the requirements someone wrote up was that multinationalizing Ruby may *not* make it any harder to work with Japanese character set than it already was, and that having to convert existing Japanese files to Unicode was simply not an option. Given all that (to my surprise) I've reached the conclusion that Unicode does *not* seem like the way to go (at least, not as part of the Ruby language proper - a good Unicode library does seem like a must). It looks like the most flexible option is to permit the creator of a string object to specify it's character width, and give the string object the capability of reporting it's own character width. Basically that's what the C++ standardisation group did - IIRC any sequence of 'character-like' objects that can be copied without causing calls to non default constructors or destructors (and especially not calls to non default copy constructors) is, technically and in principle, a standards compliant C++ string, regardless of the size in character width of those 'character-like' objects. Of course, the C++ standard also defines IO stream and stringstream libraries that can be used by both std::basic_string and std::basic_string objects... (though unfortunately not all library implementations have caught up yet). A more doable alternative to requiring such character-width independent libraries might be the use of special reader/writer classes to handle conversions to and from specific byte-widths (and perhaps also endian-ness conversions) where necessary (which is what Java does, I gather). Sincerely, Martin _________________________________________________________________ MSN Zoeken, voor duidelijke zoekresultaten! http://search.msn.nl