From: Brian Candler Date: 2011-08-10T06:38:48+09:00 Subject: Re: Is \d supposed to match Unicode Numbers? Gavin Kistner wrote in post #1015799: > Is this the correct behavior for > Ruby? Are the Oniguruma docs wrong? Am I misreading them? Do they not > apply to how Oniguruma is used within Ruby? irb(main):001:0> "0123456789٠١٢٣٤٥٦٧٨٩۰۱۲۳۴۵۶۷۸۹߀߁߂߃߄߅߆߇߈߉०१२३४५६७८९০১২৩৪৫৬৭৮৯੦੧੨…".scan(/\d/) => ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9"] irb(main):002:0> "0123456789٠١٢٣٤٥٦٧٨٩۰۱۲۳۴۵۶۷۸۹߀߁߂߃߄߅߆߇߈߉०१२३४५६७८९০১২৩৪৫৬৭৮৯੦੧੨…".scan(/[[:digit:]]/) => ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "٠", "١", "٢", "٣", "٤", "٥", "٦", "٧", "٨", "٩", "۰", "۱", "۲", "۳", "۴", "۵", "۶", "۷", "۸", "۹", "߀", "߁", "߂", "߃", "߄", "߅", "߆", "߇", "߈", "߉", "०", "१", "२", "३", "४", "५", "६", "७", "८", "९", "০", "১", "২", "৩", "৪", "৫", "৬", "৭", "৮", "৯", "੦", "੧", "੨"] irb(main):003:0> irb(main):004:0> "abcdé".scan(/\w/) => ["a", "b", "c", "d"] irb(main):005:0> "abcdé".scan(/[[:alpha:]]/) => ["a", "b", "c", "d", "é"] So I think it's intentional and consistent behaviour (for some definition of consistent): * \w and \d match only Latin letters and digits * [[:alpha:]] and [[:digit:]] match the full unicode set -- Posted via http://www.ruby-forum.com/.