From: MonkeeSage Date: 2006-09-21T12:15:05+09:00 Subject: Re: Unicode and Character Classes -- a bug? Daniel DeLorme wrote: > That is very weird indeed. It's normal that your example doesn't work, because > \xa3 is NOT valid utf8. But I would've expected it to work if you used the > correct utf8 sequence for "�" ("\xc3\xba"), except it doesn't! That shouldn't matter. He was matching the same hex escape he used in his string (viz., \xa3). It shouldn't matter whether it's unicode or just random data; the match should go through (or fail) in either case. > WTF? Can anyone explain this? Not really, because I don't understand Oniguruma (the regexp engine); I'm barely smart enough to _use_ regexps. ;) But seemingly, you can't use hex escapes in character classes, so you have to use the literal or do other things to work around it (see last two posts above). Regards, Jordan