From: Matt Armstrong Date: 2002-03-23T07:19:03+09:00 Subject: Re: Bug in Ruby Regexp? ser@germane-software.com (Sean Russell) writes: > Hi, > > Whilst trying to fix the iso-8859-1 support in Rexml, I came across > this: > > ser@valentine ~/Work/rexml $ ruby -e ' > puts /\A([^<]*)/u.match("Koln")[0]' > Koln > ser@valentine ~/Work/rexml $ ruby -e ' > puts /\A([^<]*)/u.match("K�ln")[0]' > K�ln > ser@valentine ~/Work/rexml $ ruby -e ' > puts /([^\<]*)/.match("K�ln")[0]' > K�ln Looks like the second one isn't doing what you expect. > '�' is a valid ASCII character; it's even got an ASCII value of less > than 256. To be pedantic, the iso-8859-1 '�' is not a valid ASCII character, since ASCII values only go up to 127. > The mapping for this character is the same value for unicode as it > is for iso-8859-1; Is this a bug in the unicode support in Regexp? Unicode is a superset of us-ascii only. Bytes with the high bit set have special meaning in UTF8, so you can't pass an iso-8859-1 string to routines taking UTF8 strings without some conversion. In fact, "K�ln" is not a valid UTF8 string. The byte values for each character are: K - 01001011 � - 11110110 l - 01101100 n - 01101110 < - 00111100 / - 00101111 t - 01110100 e - 01100101 s - 01110011 t - 01110100 > - 00111110 So the '�' in UTF8 means "the next 3 bytes make up one character." This explains why Ruby doesn't see the < character. However, each of the following bytes must have high bits of 10xxxxxx. Since they don't, the string isn't in UTF8 format. Ruby could legitimately detect this and throw an exception, but it doesn't. -- matt