From: Michal Suchanek Date: 2006-06-18T19:53:06+09:00 Subject: Re: Unicode roadmap? On 6/18/06, Stefan Lang wrote: > On Saturday 17 June 2006 23:55, Michal Suchanek wrote: > > On 6/17/06, Stefan Lang wrote: > [...] > > > As long as one treats a character string as a character > > > string, the internal encoding is irrelevant, and as soon as a > > > > No, it is not. > > > > First for reasons of efficiency. If an application is going to > > perform lots of slicing and poking on strings it will want some > > encoding that is suiatble for that such as UTF-32. If an > > application runs on system with little memory it will want > > space-efficient encoding (ie UTF-8 or UTF-16 for Asian languages). > > And if an appliaction runs on system that uses some legacy codepage > > it can read, write, and process all strings in that codepage. And > > in JRuby it will be useful to convert strings to UTF-16 so that the > > native Java functions can be used for manipulation. > > If you really need this level of efficiency, Ruby is probably > the wrong language anyway. Regarding JRuby: Of course each > implementation would be free to choose an internal Unicode > encoding. If somebody has enough time and motivation he can > even implement support for multiple encodings and let the user > choose at build-time. Why? It can already handle utf-8 strings or arrays of unicode codepoints. They just do not feel like strings with ruby 1.8. What I want is a glue in string class that does make them feel so. > > [...] > > > That is easy to handle with the proposed scheme: Read as much > > > as you need with the binary interface until you know the > > > encoding and then do the conversion of the byte buffer to > > > string. For file input, you can close the file when you have > > > determined the encoding and reopen it using the "normal" > > > (character oriented) interface. > > > > Why reopening or converting if you can simply tag a string that you > > had to read anyway? > > > > > Or do you mean Ruby should determine the encoding > > > automatically? IMO, that would be bad magic and error-prone. > > > > No. But if you read part of html/xml document before the encoding > > was specified there is no reason why that part hes to be converted > > or reread. You apparently got it right if you were able to > > determine the encoding from what you read. > > The conversion would be done anyway, iff a single internal > encoding was choosen and iff the encoding of the input doesn't > match the internal encoding. However, if you can choose the encoding there is no need to recode at all. You just keep the string as is, and there is a good chance the output encoding will match the input encoding. And in case you need to recode the string you got the encoding information, and the recoding can be done automatically, and only when needed. Michal