From: Rob Biedenharn Date: 2008-03-04T07:45:06+09:00 Subject: Re: Bug in regex engine ?? Must be... On Mar 3, 2008, at 2:24 PM, D. Krmpotic wrote: > Hi, > > I'm using Ruby 1.8.6, and I just discovered something rather > interesting, here is a test: > $KCODE = 'UTF8' require 'jcode' > require 'test/unit' > > class TestRegexBug < Test::Unit::TestCase > > def test_bug > > hours = "pon-čet" > > assert(hours =~ /[č]et/i) > assert(hours =~ /čet/i) > assert(hours =~ /-čet/i) > assert(hours =~ /[cč]et/i) > assert(hours =~ /-[č]et/i) > > end > > end > > As you can see, this only happens with unicode letters... (the last > test > fails).. I'm used to the fact that //i doesn't work for unicode chars > and I already know that you need two dots to match one of these.. But > this problem is different and weirder, because what triggers it is a > minus sign before the square brackets.. if you remove either the '-' > or > '[]' from the regex, it works.. > > Can you comment? > > thank you, > david Ruby is not natively aware of unicode, but you can get all these to pass if you give it the $KCOCDE hint. -Rob Rob Biedenharn http://agileconsultingllc.com Rob@AgileConsultingLLC.com