From: Bill Kelly Date: 2009-12-29T09:46:52+09:00 Subject: Re: ruby 1.9 hates you and me and the encodings we rode in on so just get used to it. Brian Candler wrote: > > I got as far as recording 200 behaviours of String in ruby 1.9 before I > gave up: > http://github.com/candlerb/string19/blob/master/string19.rb > > The solution I use is simple: stick to ruby 1.8.x. When that branch > dies, perhaps reia will be ready. If not I'll move to something else. > > IMO, both python 3 and erlang have got the right idea when it comes to > handling UTF8. Could you summarize what you feel the key difference of the python 3 / erlang approach is, compared to ruby19 ? I'm a relative newbie in dealing with character encodings, but I do recall a few lengthy discussions on this list when ruby19's M17N was being developed, where the "UTF-8 only" approaches of some other languages were deemed insufficient for various reasons. However, my understanding is that one is supposed to be able to effectively make ruby behave as a "UTF-8 only" language if one makes sure external data is transcoded to UTF-8 at I/O boundaries. I realize there may be some caveats with regard to locale, although I invoke my ruby19 scripts with -EUTF-8:UTF-8. So far, my experience with ruby19 M17N has _not_ been problematic. The only difficulties I've encountered have been when dealing with external data in some unknown encoding, where I've had to do some programmatic guesswork and finagling to make sort of a best-effort conversion of the external data to UTF-8 at the I/O boundary. But that is something I can't imagine python or erlang helping me much with either. * * * Reflecting some more, I do recall that James Gray had remarked on the difficulty of modifying one of his libraries so that it would be effectively encoding agnostic, and be able to handle data in whatever encoding was thrown at it. So from that perspective I can see how a "UTF-8 only" approach at the language level should simplify things. But from my current perspective as an application developer who is taking the approach of ensuring all data read into my program is converted to UTF-8, I'm wondering if my experience is essentially similar to what it would be in a "UTF-8 only" language. Regards, Bill