From: Charles O Nutter Date: 2006-06-26T11:37:45+09:00 Subject: Re: Unicode roadmap? ------=_Part_40099_16859329.1151289461086 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: 7bit Content-Disposition: inline One clarification I'd like to add to this: I'm not saying that a ByteArray needs to be added, but if you're going to treat String as a ByteArray, then perhaps there should be another type for character vectors? Perhaps through some logic (perhaps the fact that this is the "way it is" in Ruby 1.8) String does == ByteArray. If I could play devil's advocate for a moment, maybe the new, fancy m17n String, however it's implemented, should be a different class? String == ByteArray in form and function CharString == a string of characters with some particular encoding, character logic, and so on Perhaps even CharString < String, so it retains byte-level read/write operations. There's another obvious advantage here...APIs that currently return a byte array String will continue to do so, as they work in Ruby 1.8. CharString could also be implemented today for Ruby 1.8, providing an encoding and character-aware String implementation for applications that need it. My only point about the dichotomy between and is that at some level, they imply different behaviors, different APIs, different interfaces. Perhaps the answer is not to change existing Ruby code to use a m17n String while trying to retain byte array capabilities in the same time...but maybe it's worth considering that the new behavior warrants a separate type? String.to_cs(:utf8) => CharString String retains current interface and semantics CharString gains the [n] => character or single-char string rather than int, etc. I know you (matz) want to break as much as possible with the 2.0 release, but I still don't see the advantage of marrying the "byte array string" and "char string" types in the same class when separate types and behaviors would be more logical and break far less. On 6/25/06, Charles O Nutter wrote: > > On 6/25/06, Austin Ziegler wrote: > > > > Ruby does not need a String with an internal representation in Unicode; > > Ruby does not need a separate byte vector. An unencoded string can be > > treated as a byte vector with no problems; if it is determined to have > > textual meaning, it can be tagged with an encoding very simply and from > > that point be treated as a meaningful string. There are times when the > > encoding is *not* best treated in Unicode, especially if there are > > potential conversion errors. > > > > When is a ByteArray not a ByteArray? When is a String not a String? Is it > correct to mingle the two concepts perpetually, when they each have fairly > specific definitions? My problem with continuing to treat String as a byte > vector is that it forces two somewhat incompatible concepts on the same > class and the same methods. If you can use a String as both a byte vector > and as a sequence of characters by calling the same methods, then setting or > clearing encoding suddenly has the side-effect of changing how elements of > String are to be treated. If you are providing separate methods for working > with bytes as opposed to working with characters, then you are already > splitting the two concepts. > > (As an aside, does it make sense that I read from a binary file into a > String? Can I reliably assume that binary content in a String should be > logically manipulable as text strings are? Should my binary String work > anywhere and everywhere a text-based String does? I would think that binary > content neither walks nor quacks like a String.) > > By your definition, a String can be treated as a ByteArray so long as its > internal string does not have an encoding. What do I use if I want to have > an encoding and still use byte vector semantics? > > It is appropriate that a String is no longer usable as a ByteArray as a > result of changing some state? If there exists any state where String cannot > be logically treated as a byte array, then String != ByteArray in the > general case either. The encoding of a String's internal representation > should not dictate the outward behavior of the String. > > If, however, you completely separate the two concepts, there's no > dichotomy. In that case, a String deals with characters, and you do not have > guarantees about byte-boundaries or indexed elements. You only have > guarantees about characters, as it should be. Simultaneously, ByteArray > would allow you to always work with a vector (array) of bytes, regardless of > what those bytes contain. > > I'll end it off saying this: I think it's a no-brainer that for dealing > with streams of bytes, there should be a non-string byte vector class. If > folks are insistent on keeping them the same class, you can't logically > continue to call it a String and have it fulfill the dual purposes of byte > vector and character vector at the same time. If you plan to provide methods > for supporting both behaviors, you're putting two distinct behaviors into > the same type. > > I understand the unwillingness to move away from String as a byte vector, > but with multibyte support coming you really can't have String == ByteArray > without causing problems somewhere. They simply don't have the same > behavior, and trying to pretend they do is asking for trouble. > > > -- > Charles Oliver Nutter @ headius.blogspot.com > JRuby Developer @ www.jruby.org > > Application Architect @ www.ventera.com > -- Charles Oliver Nutter @ headius.blogspot.com JRuby Developer @ www.jruby.org Application Architect @ www.ventera.com ------=_Part_40099_16859329.1151289461086--