From: Mark Hubbart Date: 2005-05-27T12:54:30+09:00 Subject: Re: Using string.slice for unicode chars On 5/26/05, Ilya V. Sabanin wrote:> Hi,> > I try to cut unicode string with "slice" method but looks like Ruby> unable to do this.> Of course I requiring 'jcode':> > $KCODE = 'u'> require 'jcode'> > So, Ruby unable to cut utf-8 strings? Not yet. I think it's slated for 2.0 (or 1.9 first) Some possible workarounds: ruby 1.8.2 (2004-12-25) on powerpc-darwin8.0.0Welcome to Interactive Ruby us = 'Ṧȶȑïŋɠ' ==>"Ṧȶȑïŋɠ" us.scan(/./) ==>["Ṧ", "ȶ", "ȑ", "ï", "ŋ", "ɠ"] us.unpack('U*') ==>[7782, 566, 529, 239, 331, 608] regexen will abide by the kcode, or you can unpack the utf-8 string asan array of integer code points. It shouldn't be too hard to wrap the regexp idea into a u_slicemethod. here's an untested lightwieght version: class String def u_slice(index, size = 1) self[/.{#{index}}(.{#{size}})/, 1] end def u_slice!(index, size = 1) str = self[/.{#{index}}(.{#{size}})/, 1] self[/.{#{index}}(.{#{size}})/, 1] = "" str end end