From: Dae San Hwang Date: 2006-06-15T00:47:12+09:00 Subject: Re: A plan for another unicode string hack On Jun 14, 2006, at 11:56 PM, Logan Capaldo wrote: > > On Jun 14, 2006, at 10:47 AM, Dae San Hwang wrote: > >> The reason I'm differentiating between 'size' and 'length' is >> because some libraries (like rails) depend on them returning the >> byte size of the string. Maybe we can establish a customs that >> 'size' for byte size and 'length' for the number of characters. >> Same reasoning goes for '[]' and 'slice'. > > I like these very much. Although the choice between [] and slice > seem arbitrary (i.e. you could have swapped their meanings and it > would have made just as much sense). #size vs. #length is perfect. > and #[] being a Fixnum when their was no encoding but a character > when there is is equally brilliant. I salute you sir! > Thanks for the kind words. The reason I picked [] for encoding aware method is because String# [index] will be used to extract the letter and not the byte in Ruby 2.0 as mentioned in http://redhanded.hobix.com/inspect/ futurismUnicodeInRuby.html so that "abc"[0] returns "a" instead of fixnum 97 A way to get a Nth byte of a byte buffer is probably still necessary and String#slice seems to be the logical one, I thought. Dae San Hwang daesan@gmail.com