From: Frank Mitchell Date: 2002-06-29T13:07:25+09:00 Subject: Re: Andy Roonie is perhaps excessively optimistic Benjamin Peterson wrote: > If Ruby were my own project, I might do the following: > > 1 -- Bite the Unicode bullet and accept that despite > the legitimate concerns of many Japanese, it is the > standard and it works well enough to get things done. > I wish TRON had won. It didn't win. > 2 -- Use wide characters in Ruby internally. Forget > surrogates, like every other implementation does. > UCS2 characters are fast (always the same length) and > pretty near standard. A change like this is worth > having to recompile things for. > 3 -- Plug in rxpp regular expressions to replace the > narrow-character gnu regex file we have. > 4 -- Isolate IO routines (including console IO) to > provide a layer for translating encodings. There > could be more than one layer (I would want a > windows-specific one, but to start with you could just > put in a dumb 'squashing' of the internal UCS2 to > ASCII). All translations would be between UCS2 and > the currently active IO encoding. Java programmers will tell you that converting Unicode to a native encoding takes up a surprisingly large amount of time. Reading a string from a file, doing a trivial substitution, and writing it to another file does an unnecessary amount of work. Granted, nobody expects a Ruby script to be blindingly fast, but other threads in this newsgroup are complaining about I/O being slow. Maybe this has been suggested already but, since Ruby is object-oriented, I'd vote for two (or more) virtually indistinguishable String classes, one for Unicode strings, one for single-byte strings. Perhaps byte strings could have an "encoding" attribute (a Symbol) to make converting from one representation to another automatic. Maybe you'd also need a distinction between getting the Nth byte, and getting the Nth character (always converted to a Unicode character.) Note that Python has two string types which are virtually indistinguishable. Every string function "does the right thing" whether it operates on a byte string or a wide string, as far as I know. (I haven't tried the regexp functions yet.) -- Frank Mitchell (frankm@bayarea.net) Please avoid sending me Word or PowerPoint attachments. See http://www.fsf.org/philosophy/no-word-attachments.html