From: brabuhr@... Date: 2010-06-18T01:04:28+09:00 Subject: Re: nothing new in ruby_core for four days ? >>> But >>> 1.9.2 implies (or is it 1.9.x ?) - alas poor Yorick! -  changing >>> Regexp's : \d, \s,\w and the like work no more with utf-8 encodings! >> >> Care to share an example, or point to a blog with more details? I have >> an older 1.9.2 lying around, it doesn't seem to behave how you describe. > > http://svn.ruby-lang.org/repos/ruby/tags/v1_9_2_preview3/NEWS > * \d, \s, and \w are now ASCII only; use POSIX bracket classes and \p{} for >  Unicode semantics > > http://redmine.ruby-lang.org/issues/show/3376 > The rationale is that many people expects \d and \s to match ASCII only. http://blade.nagaokaut.ac.jp/cgi-bin/scat.rb/ruby/ruby-core/30543 For your convenience, I tell you two well-known incompatibilities: - \d, \s and \w in Regexp does match *only* ASCII characters. Especially, \w does NOT match with accented character, umlaut, etc. In other words, it behaves the same as 1.8. # coding: UTF-8 p(/\w/ =~ "a") #=> nil If you want Unicode sensitive character class, you can use \p{Digit} instead of \d, \p{Space} instead of \s, and \p{Word} or [\w\P{ASCII}] instead of \w.