From: Austin Ziegler Date: 2006-06-29T03:43:54+09:00 Subject: Re: Unicode roadmap? On 6/28/06, Juergen Strobel wrote: > Any additional complexity here should be offset later, when doing > operations on the read data as appropriate for its type. It won't be. All of the complexity of the m17n String will be inside of the String, not exposed (by default) to the user. Stop thinking of the encoding of a String as something that makes the String a unique object; instead it is a lens that gives meaning to the bytes of the String. > Of course, the first line should raise an exception if file.txt is not > utf8 encoded, The internal format of String is not going to be Unicode by default. Matz has already said that. I happen to agree with him. > this saves extra complexity down the line, and is a real difference > between the two. I imagine Bytevector would be implemented with > maximum performance and space efficiency in mind, while String is a > higher level class streamlined for easy of use. These two items are not mutually exclusive. Think a little more about humane design and you'll see that two wholly separate classes require a lot more than what you're assuming and would end up in programmers making even dumber assumptions than they do today, because they'd think they're "protected" during IO because they're getting a String. This is not a safe assumption. Ever. The separate byte vector class is needlessly complex and solves exactly nothing that isn't already solved in a better way. -austin -- Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ * austin@halostatue.ca * http://www.halostatue.ca/feed/ * austin@zieglers.ca