From: Curt Sampson Date: 2002-08-05T08:17:16+09:00 Subject: Re: Unicode in Ruby? On Mon, 5 Aug 2002, Jan Witt wrote: > So, ultimately, as computers and networks are increasing performance > daily, why can we not concentrate on functionality? Once we get > language across, we will be able to get meaning across. The question is really about where you put the functionality. Should all this be at the lowest-level of string handling (your basic string class), or should some of it be up at higher levels? If you look at some common applications, such as dealing with small strings of data that move between a user's web browser and a service provider's database (login names, passwords, addresses, etc.), there's very little processing to be done. Often you don't even care what language or anything else the data are in; you just need to know the character set encoding (so you don't mix encodings on output). The web browser takes care of most or all of your input and display issues, including even wrapping text. (And that's a big, big issue right there; Japanese and English use quite different text wrapping algorithms, for example.) Unicode is designed to let you leave the lower levels quite simple, and use upper-level code to deal with more complexity only if you really need to. Moving this kind of stuff into lower level code leads to several problems: 1. It's not efficient to have lower level code do things that you don't need it to do. This is not a killer problem, I agree, but is still a concern. 2. Having the lower-level code doing this stuff can actually force unnecessary complexity on to the upper-level stuff using that code. 3. If the lower-level stuff is broken in some way, or doesn't do all that you need, you still end up having to have further processing in the upper layers anyway. To my mind, reason number three is the real killer: text processing varies in complexity and can get very, very, very complex. And for those applications, there's a good chance that the lower levels are actually going to do things wrong in some cases, in which case it makes life far more complex for the programmer than if it had just stayed out of the way in the first place. Take even the simple idea of iterating over the characters in a string, for example. As soon as you bring in left-to-right languages, such as Hebrew, how does this work? Do you iterate in the order the text is read? What happens when the direction of reading changes in the middle? (Such as when you mix Hebrew and arabic numerals?) What if it's a multi-line string? Investigate the issues and work out how you might deal with this, and you'll quickly realize that you really want all this complexity in a class separate from your basic string. cjs -- Curt Sampson +81 90 7737 2974 http://www.netbsd.org Don't you know, in this new Dark Age, we're all light. --XTC