From: Mauricio Fernandez Date: 2006-10-17T06:04:25+09:00 Subject: Re: Edit Distance at Wikipedia On Mon, Oct 16, 2006 at 10:11:51PM +0900, Minkoo wrote: > On 10/16/06, Mauricio Fernandez wrote: > >It's simply the string you're comparing against; unpack('U*') just turns > >the UTF-8 characters into unsigned integers: > > > > class String > > def levenshtein(comparator) > > a, b = self.unpack('U*'), comparator.unpack('U*') > > b # => [102, 111, > >111, 98, 97, 114] [...] > > > > "foo".levenshtein("foobar") # => 3 > > > I'm afraid that I'm not used to character encodings. Does Ruby use UTF-8 by > default? Ruby's Strings as of 1.8 are just a sequence of bytes. If a String happens to hold a UTF-8 sequence, some operations will work properly when you set $KCODE="u", e.g. s = "\303\241" # "�" == acute s.scan(/\w/) # => [] $KCODE="u" s.scan(/\w/) # => ["�"] Some other methods are covered by jcode.rb. > In other words, suppose that I've launched irb and fired > "foo".levenshtein("foobar"). > In that case, is the string "foo" encoded as utf-8? This depends on your locale; Ruby's Strings do not know about encodings in 1.8 (they will very soon in 1.9 according to matz' latest messages in ruby-core; search the ruby-talk archives for extensive discussions of the upcoming M17N). > Do I always have to unpack the string like the code shown above? In this particular case, it depends on two things: * the encoding you're using * whether the Levenshtein should be relative to characters or bytes In general, it's up to you to keep track of the encodings and handle Strings appropriately. -- Mauricio Fernandez - http://eigenclass.org - singular Ruby