From: Stefan Lang Date: 2006-06-18T09:15:59+09:00 Subject: Re: Unicode roadmap? On Saturday 17 June 2006 23:55, Michal Suchanek wrote: > On 6/17/06, Stefan Lang wrote: [...] > > As long as one treats a character string as a character > > string, the internal encoding is irrelevant, and as soon as a > > No, it is not. > > First for reasons of efficiency. If an application is going to > perform lots of slicing and poking on strings it will want some > encoding that is suiatble for that such as UTF-32. If an > application runs on system with little memory it will want > space-efficient encoding (ie UTF-8 or UTF-16 for Asian languages). > And if an appliaction runs on system that uses some legacy codepage > it can read, write, and process all strings in that codepage. And > in JRuby it will be useful to convert strings to UTF-16 so that the > native Java functions can be used for manipulation. If you really need this level of efficiency, Ruby is probably the wrong language anyway. Regarding JRuby: Of course each implementation would be free to choose an internal Unicode encoding. If somebody has enough time and motivation he can even implement support for multiple encodings and let the user choose at build-time. [...] > > That is easy to handle with the proposed scheme: Read as much > > as you need with the binary interface until you know the > > encoding and then do the conversion of the byte buffer to > > string. For file input, you can close the file when you have > > determined the encoding and reopen it using the "normal" > > (character oriented) interface. > > Why reopening or converting if you can simply tag a string that you > had to read anyway? > > > Or do you mean Ruby should determine the encoding > > automatically? IMO, that would be bad magic and error-prone. > > No. But if you read part of html/xml document before the encoding > was specified there is no reason why that part hes to be converted > or reread. You apparently got it right if you were able to > determine the encoding from what you read. The conversion would be done anyway, iff a single internal encoding was choosen and iff the encoding of the input doesn't match the internal encoding. > > [...] > > > > > > If the strings are represented as a sequence of Unicode > > > > codepoints, it is possible for external libraries to > > > > implement more advanced Unicode operations. > > > > > > This would be true regardless of the encoding. > > > > But a conversion from [insert arbitrary encoding here] to > > unicode codepoints would be needed. > > That will be needed anyway. You cannot expect all libraries to use > the arbitrary encoding you chose for Ruby strings. I assume you mean C libraries here. -- Stefan