From: Ammar Ali Date: 2010-11-10T03:09:59+09:00 Subject: Re: String#split regex \W on non-ASCII text On Tue, Nov 9, 2010 at 7:44 PM, Fritz Anderson wrote: > Ruby 1.9.2-p0, built from the tarball on Mac OS X 10.6.4 with the Xcode > 3.2.5 tools. > > Consider the following string: > STR = "sarà la cortesia del gran Lombardo" > > Ruby (through irb) correctly identifies this as a Unicode string; the > first word (in case something swallows it on the way to your screen) > ends with an a-grave. > > I'd like to naïvely split this line into words. The obvious way to do > this is: > > words = STR.split /\W+/ > # adding the u qualifier to the regexp doesn't matter > > words becomes > => ["sar", "la", "cortesia", "del", "gran", "Lombardo"] > Use the unicode property for separators. >> words = STR.split /\p{Z}/ => ["sarà", "la", "cortesia", "del", "gran", "Lombardo"] Regards, Ammar words = STR.split /\p{Z}/