From: Hermann Ulrich Date: 2010-12-06T03:52:16+09:00 Subject: Re: String#split regex \W on non-ASCII text Am 09.11.2010 um 19:09 schrieb Ammar Ali: > On Tue, Nov 9, 2010 at 7:44 PM, Fritz Anderson wrote: >> Ruby 1.9.2-p0, built from the tarball on Mac OS X 10.6.4 with the Xcode >> 3.2.5 tools. >> >> Consider the following string: >> STR = "sarà la cortesia del gran Lombardo" >> >> Ruby (through irb) correctly identifies this as a Unicode string; the >> first word (in case something swallows it on the way to your screen) >> ends with an a-grave. >> >> I'd like to naïvely split this line into words. The obvious way to do >> this is: >> >> words = STR.split /\W+/ >> # adding the u qualifier to the regexp doesn't matter >> >> words becomes >> => ["sar", "la", "cortesia", "del", "gran", "Lombardo"] >> > > Use the unicode property for separators. > >>> words = STR.split /\p{Z}/ > => ["sarà", "la", "cortesia", "del", "gran", "Lombardo"] > > Regards, > Ammar > words = STR.split /\p{Z}/ > >