From: Charles O Nutter Date: 2006-06-18T14:00:48+09:00 Subject: Re: Unicode roadmap? ------=_Part_49369_26659330.1150606845603 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: 7bit Content-Disposition: inline I'll chime back in with my not-so-expert opinion, so it's known where I stand. Take it for whatever it's worth. - I almost entirely agree with Juergen's longer post on what unicode support should look like in 2.0. I won't go into the details of what I disagree with because I'm a little squishy in those areas. - I believe that supporting encoding-tagged strings would be a horrible, horrible mess for both Ruby VM/interpreter implementers and extension implementers while not adding any serious benefits for Ruby the language. When it comes down to it, you're going to have string A using encoding X and string B using encoding Y and in order to work with them both together you'll have to find some common ground. Settle on common ground early or you pay the price to do it EVERY time you work with strings later. - I have no intention to ever write a C extension for Ruby. I know many out there do. However, I think the important thing about Ruby is Ruby, and making the language bend over backwards to make life easier for C hackers is absurd. Making unicode support needlessly complex in Ruby (the language) only ends up hurting its usability. I for one would not want to sacrifice the beauty and simplicity of Ruby solely to apease the C community. Flame on if you will, but The Ruby Way should rule here. - In the end, I should not have to care what encoding strings use internally unless I absolutely have to know. Every time questions come up about unicode support in Java, I have to look it up...UTF-8? UTF-16? UCS-2? I rarely need to know this information, and I rarely remember it. That's exactly the point. Make the one internal encoding whatever is deemed most flexible, most performant, and above all *most global*. Nobody writing Ruby code should have to care. - I so rarely work with Strings on a character-by-character basis, and when I do all I should have to say is get_character and know that what I have represents a full and complete character representation. If you're dealing with bytes, call it what it is--the aforementioned ByteBuffer. Ruby needs to support the concepts of Strings and ByteBuffers independently. I think it all comes back to a simple question: Which method of supporting unicode would feel the most "Ruby"? Which one is DRY and KISS and all the other lovely acronyms this community holds so dear? Figure that out, and there's your answer. I'd be willing to bet it's not every-string-can-encode-differently, because I don't see how that would ever help me write better Ruby code...and improving Ruby is the point of all this, right? -- Charles Oliver Nutter @ headius.blogspot.com JRuby Developer @ jruby.sourceforge.net Application Architect @ www.ventera.com ------=_Part_49369_26659330.1150606845603--