From: Dossy Date: 2002-07-07T22:43:57+09:00 Subject: Re: is there a better string.each? On 2002.07.07, Tobias Reif wrote: > Dossy wrote: > > > >So, perhaps an easy way to avoid backward-compatibility breakage > >and to ease into m17n, we simply define String#characters which > >returns an Array of Character -- which can be either single-byte, > >or multi-byte, depending on the appropriate encoding. Then, > >Character#each would Do The Right Thing and iterate over the > >array, one Character at a time. > > > String#characters would be cool. > > But I don't understand the rest: what does it return? > I'm confused about "returns an Array of Character" and "Character#each". > An object of class Character represents one character; else it would > be "Characters", which is more of a string. If you want an array of > characters, do > > 'foo'.split '' > # ["f", "o", "o"] > > anywas; it would probably pass the chars to the black *and* return an array. This fails when the string is encoded in a multi-byte encoding (such as Japanese, Chinese, etc.) unless String#split is modified in the core to split on actual logical "characters" (which might be one or more bytes, depending on the encoding). Until that time, I was suggesting that _today_ in Ruby 1.6 we can support multiple encodings by jamming a String#characters in that did understand encodings. I'd even go so far as to say we introduce String#encoding= and String#encoding to be able to tell String class what encoding a specific instance of a String is. I was thinking something along the lines of: 'foo'.characters.type # => Array 'foo'.characters # => ['f', 'o', 'o'] 'foo'.characters[0].type # => Character The idea here though is that String#characters would handle strings that were in multi-byte encodings and Do The Right Thing ... > I think String#characters should pass objects of class Character to the > block. > > 'foo'.chars do |char| > puts char > end > # f > # o > # o That's limiting. What if you wanted an Array of the Character objects? You'd have to do something like: charArray = [] 'foo'.chars { |char| charArray << char } Foul. Having it return an Array makes it so much easier to handle. > The increase in convenience would not be too dramatic in some cases: > > irb(main):005:0> 'foo'.split('').each do |char| puts char, char.class end > f > String > o > String > o > String > ["f", "o", "o"] > > > 'foo'.chars do |char| puts char, char.class end > f > Character > o > Character > o > Character > ["f", "o", "o"] > > ... but a class Character might be a good thing for i18n. Right. This is all about driving towards i18n and m17n today, without having to go (1) break backwards compatibility or (2) change the core. As its been pointed out before, though: we'll have to wait for the new Regex code for regexes to Do The Right Thing with regard to multi-byte encoded strings and regexes. -- Dossy -- Dossy Shiobara mail: dossy@panoptic.com Panoptic Computer Network web: http://www.panoptic.com/ "He realized the fastest way to change is to laugh at your own folly -- then you can let go and quickly move on." (p. 70)