From: Brian Candler Date: 2009-08-12T18:18:00+09:00 Subject: Re: 1.9 String and M17N documentation BTW, I find James's writeup of what he had to do to the CSV library (*) enlightening. Even ruby 1.9 won't match an ASCII regexp like /,/ against a wide encoding, so he had to generate new regexps dynamically at runtime. Now, I think that's a good thing, optimising the regexps to match the incoming data stream efficiently. But I also observe that this would have worked just fine if the encoding were a property of the regexp only - which is the approach 1.8 takes to regexps. What I mean is, once you have decided to build a "UTF-16LE" regexp, say, you can just match it against a stream of bytes. Making every single String also have an encoding property only gives more opportunities for Ruby to raise exceptions. Some may argue this is Ruby "protecting" you from doing something silly, but if I'm working with string literals or binary data returned from a library, whose encoding may or may not have been set to ASCII-8BIT, then I don't want this "protection". Rather, I need protecting against ruby 1.9. There is only one case I can see where having the encoding be a property of the String itself is useful: selecting individual characters by index. e.g. if str.size > 50 str = str[0,47] + "..." end There's a huge amount of language pain introduced just for that. Regards, Brian. (*) http://blog.grayproductions.net/articles/what_ruby_19_gives_us -- Posted via http://www.ruby-forum.com/.