From: MonkeeSage Date: 2007-12-03T10:50:00+09:00 Subject: Re: Unicode in Regex On Dec 2, 2:35 pm, Greg Willits wrote: > Greg Willits wrote: > > I'm expecting a validate_format_of with a regex like this > > /^[a-zA-Z\xC0-\xD6\xD9-\xF6\xF9-\xFF\.\'\-\ ]*?$/ > > to allow many of the normal characters like ö é å to be submitted via > > web form. However, the extended characters are being rejected. > > So, I've been pounding the web for info on UTF8 in Ruby and Rails the > past couple days to concoct some validations that allow UTF8 > characters. I have discovered that I can get a little further by doing > the > following: > - declaring $KCODE = 'UTF8' > - adding /u to regex expressions. > > The only thing not working now is the ability to define a range of \x > characters in a regex. > > So, this /^[a-zA-Z\xE4]*?&/u will validate that a string is allowed > to have an ä in it. Perfect. > > But... this fails /^[a-zA-Z\xE4-\xE6]*?&/u > > But... this works /^[a-zA-Z\xE4\xE5\xE6]*?&/u > > I've boiled the experiments down to realizing I can't define a range > with \x > > Is this just one of those things that just doesn't work yet WRT Ruby/ > Rails/UTF8, or is there another syntax? I've scoured all the regex > docs I can find, and they seem to indicate a range should work. > > For now, I just have all the characters I want included < \xFF listed > individually. > > utf_accents = '\xC0\xC1\xC2\.......' > > Is_person_name = /^[a-zA-Z#{utf_accents}\.\'\-\ ]*?$/u > > But I'd like to solve the range notation if I can. > > -- > def gw > acts_as_n00b > writes_at(www.railsdev.ws) > end > -- > Posted viahttp://www.ruby-forum.com/. This seems to work... $KCODE = "UTF8" p /^[a-zA-Z\xC0-\xD6\xD9-\xF6\xF9-\xFF\.\'\-\ ]*?/u =~ "Jäsp...it works" # => 0 However, it looks to me like it would be more robust to use a slightly modified version of UTF8REGEX (found in the link Jimmy posted above)... UTF8REGEX = /\A(?: [a-zA-Z\.\-\'\ ] | [\xC2-\xDF][\x80-\xBF] | \xE0[\xA0-\xBF][\x80-\xBF] | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2} | \xED[\x80-\x9F][\x80-\xBF] | \xF0[\x90-\xBF][\x80-\xBF]{2} | [\xF1-\xF3][\x80-\xBF]{3} | \xF4[\x80-\x8F][\x80-\xBF]{2} )*\z/mnx p UTF8REGEX =~ "Jäsp...it works here too" # => 0 Look at the link to see the explanation of the alternations. Regards, Jordan