From: Brian Candler Date: 2010-06-21T19:10:08+09:00 Subject: Re: String comparison. Why does Ruby consider this true? Josh Cheek wrote: > Well, this used to be easy to show, but apparently since ascii has been > abandoned, and I don't know unicode, I have to resort to hacky things > like > this to explain it. > > > $chars = (1..128).inject(Hash.new) { |chars,num| chars[num.chr] = num ; > chars } > > def to_number_array(str) > str.split(//).map { |char| $chars[char] } > end > > to_number_array 'Xeo' # => [88, 101, 111] > to_number_array 'xeo' # => [120, 101, 111] > to_number_array 'ball' # => [98, 97, 108, 108] > to_number_array 'ABC' # => [65, 66, 67] > to_number_array 'abc' # => [97, 98, 99] Except that this is irrelevant, because even ruby 1.9 does not compare strings by codepoints. It compares them byte-by-byte using memcmp. See rb_str_cmp_m() and rb_str_cmp() in string.c It's a designed-in side-effect of UTF-8 encoding that higher codepoints sort after lower ones. There is a table at http://en.wikipedia.org/wiki/UTF-8 under "Description" which illustrates this. However this does not work for other encodings. Try this for size: >> s1 = 97.chr("UTF-8") => "a" >> s2 = 257.chr("UTF-8") => "ā" >> s1 < s2 => true >> s1 = 97.chr("UTF-16LE") => "a\x00" >> s2 = 257.chr("UTF-16LE") => "\x01\x01" >> s1 < s2 => false Yes: that's the same two unicode codepoints, but sorting in different order. For encodings like UTF-16LE, where the least-significant byte comes before the most-significant byte, you get an almost arbitrary ordering. Proviso: I tested this with ruby 1.9.2dev (2009-07-18 trunk 24186) [i686-linux] ruby 1.9.x string encoding rules are (a) undocumented, and (b) subject to arbitrary changes between patchlevels, hence YMMV. -- Posted via http://www.ruby-forum.com/.