From: Charles O Nutter Date: 2006-06-29T05:43:32+09:00 Subject: Re: Unicode roadmap? ------=_Part_102796_11131418.1151527407519 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: 7bit Content-Disposition: inline I'll give a little ground on a few points. Perhaps I had a dream that adjusted my perspectives a bit. - String == ByteArray is reasonable if String is considered to be a "ByteChunker". The byte-chunking logic holds true for both binary string and encoded string models; what's parametric about it is the size of the chunks. While I still believe that a general-purpose, high-performance ByteArray would be useful (perhaps preferable) for many operations, I will concede that ByteArray + ChunkSizer (encoding) := m17n String. I won't say I'm sold on the ByteChunker pattern, but I think it will be easier to accept and discuss m17n Strings from a ByteChunker perspective. It also may be fair to say that m17n String provides a "view" into the underlying byte array, which could in the raw case be a wholly-transparent view. - If the intent is to provide a String that supports all encodings universally, I will concede that the m17n String is probably the only way to do it. As far as I know, there's no one character encoding, code page, or character set that can encompass all other encodings without fail. Unicode does, despite what detractors may say, make a truly gallant attempt to achieve that impossible goal, and it deserves the 90+% of humanity that use Unicode or Unicode-encodable character sets exclusively. But if at the end of the day Ruby really needs a kitchen-sink approach to character encoding, Unicode will not fit that requirement. So, a short glossary: String == ByteChunker chunk == character == n bytes in a specific order The first item above brings out a few discussion points 1. String provides an interface for managing a collection of ByteChunks. The sizing and nature of this "chunking" is primarily based on character encoding logic. I'll refer to String and ByteChunker interchangeably from here on out. 2. Indexed operations act upon chunks, not bytes. It may be the case that for some encodings, sizeof(chunk) == sizeof(byte). No assumptions should be made about chunk size. 3. Altering String semantics from "byte ops always" to "chunk ops always" also implies that chunked operations should not be generally purposed toward byte-level operations, since there is no explicit guarantee you'll work with byte-sized (har har) chunks 4. Therefore it should be mandatory and acceptable under the supposed ByteChunker contract to provide a minimal set of explicitly byte-sized operations, since the purpose of chunking is to provide a way of consuming and digesting bytes. It would not be useful or recommended to completely hide those raw bytes under any circumstances, since byte-level operations will always be valid on a ByteChunker in the absence of a more specific ByteArray type. 5. Byte-sized operations should be STRONGLY ENCOURAGED for byte-level work over chunk-sized operations due to the changing size and nature of chunks. This would mean that [0..5] should never be used instead of byte(0..5) for retrieving the first five bytes in a ByteChunker 6. Methods on other classes whose purpose is to manipulate character data (chunks) logically should never be assumed to work with byte-sized chunks only (regex and friends) A common theme here is that ByteChunker does have a set of logical semantics, and m17n Strings as planned appear to be ByteChunkers. This seems like a reasonable abstraction to me, though it does expose implementation details many of us would prefer to keep hidden (namely that we're chunking bytes, when a consumer shouldn't need to know what chunks are composed of). If we can reasonably attempt to define the ByteChunker semantics, we can see where the holes are. I could do without a separate ByteArray if the m17n String provided explicit byte-sized operations. The dual-purposing of String ops for both bytes and chunks is very worrisome since it's bound to happen that chunk operations get incorrectly used for byte operations when sizeof(chunk) != sizeof(byte). For byte-sensitive cases, rather than saying "I know that my String is in encoding X in which all chunks are byte-sized" it would be far safer (and better encapsulation of String state) to say "I know I need to work with bytes all the time". The byte-sized operations then follow. (Granted, there's no way to force people *ahem Austin* to use the "byte-safe" methods, but it feels like a really good best-practice to me). A caveat to all this is that ByteChunker semantics are inherently more complex than CharacterSequence semantics, and so the proposed m17n String *is* a more complicated solution than using a single internal encoding. I'm also not convinced that ByteChunker's semantics are simpler than separate CharacterSequence and ByteArray semantics, they they may be more "Ruby." ByteChunker itself is a more useful general-purpose entity in itself than CharacterSequence or ByteArray alone. complexity(ByteChunker) > (complexity(ByteArray) or complexity(CharacterSequence)) complexity(ByteArray and CharacterSequence) maybe > complexity(ByteChunker) generality(ByteChunker) > (generality(ByteArray) or generality(CharacterSequence)) I think it's still a valid question whether there's not a happy medium somewhere that would make life easiest for the folks using unicode. Rubyists are fond of saying that Ruby makes easy problems easy and hard problems possible. I would argue that unicode support should be the "easy problem" that's easy and that support for incompatible encodings--worldwide--should be the "hard problem" that's possible. Any plans for m17n that make unicode harder to work with in Ruby than in comparable languages could prove fatal. -- Charles Oliver Nutter @ headius.blogspot.com JRuby Developer @ www.jruby.org Application Architect @ www.ventera.com ------=_Part_102796_11131418.1151527407519--