From: Gavin Kistner Date: 2011-08-10T05:28:04+09:00 Subject: Is \d supposed to match Unicode Numbers? --Boundary_(ID_gHyJdGyCH+UMzqa1eNhBXQ) Content-type: text/plain; charset=utf-8; format=flowed Content-transfer-encoding: quoted-printable I posted this as a question here:=0Ahttp://stackoverflow.com/questions/699= 8713/scanning-for-unicode-numbers-in-a-string-with-d=0A=0ASummarized:=0ATh= e Oniguruma docs[1] seem to say that \d is supposed to match the Unicode "= Decimal_Number" category. However, in Ruby 1.9.1 and 1.9.2 it only matches= Latin 0-9 characters. Is this the correct behavior for Ruby? Are the Onig= uruma docs wrong? Am I misreading them? Do they not apply to how Oniguruma= is used within Ruby?=0A=0ATest program:=0A=0A#encoding: utf-8=0Arequire '= open-uri'=0Ahtml =3D open("http://www.fileformat.info/info/unicode/categor= y/Nd/list.htm").read=0Adigits =3D html.scan(/U\+([\da-f]{4})/i).flatten.ma= p{ |s| s.to_i(16) }.pack('U*')=0A=0Aputs digits.encoding, digits=0A#=3D> U= TF-8=0A#=3D> 0123456789=D9=A0=D9=A1=D9=A2=D9=A3=D9=A4=D9=A5=D9=A6=D9=A7=D9= =A8=D9=A9=DB=B0=DB=B1=DB=B2=DB=B3=DB=B4=DB=B5=DB=B6=DB=B7=DB=B8=DB=B9=DF=80= =DF=81=DF=82=DF=83=DF=84=DF=85=DF=86=DF=87=DF=88=DF=89=E0=A5=A6=E0=A5=A7=E0= =A5=A8=E0=A5=A9=E0=A5=AA=E0=A5=AB=E0=A5=AC=E0=A5=AD=E0=A5=AE=E0=A5=AF=E0=A7= =A6=E0=A7=A7=E0=A7=A8=E0=A7=A9=E0=A7=AA=E0=A7=AB=E0=A7=AC=E0=A7=AD=E0=A7=AE= =E0=A7=AF=E0=A9=A6=E0=A9=A7=E0=A9=A8=E2=80=A6=0A=0Ap RUBY_DESCRIPTION, dig= its.scan(/\d/)=0A#=3D> "ruby 1.9.2p180 (2011-02-18) [i386-mingw32]"=0A#=3D= > ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9"]=0A=0AFeel free to dis= cuss here, or answer on Stack Overflow if you have a solid answer and want= the rep :)=0A=0A=0A[1]=C2=A0http://www.geocities.jp/kosako3/oniguruma/doc= /RE.txt=0A=0A--=0A(-, /\ \/ / /\/= --Boundary_(ID_gHyJdGyCH+UMzqa1eNhBXQ) Content-type: multipart/related; boundary="Boundary_(ID_59lsIthJotE8yodTDJkLtA)"; type="text/html" --Boundary_(ID_59lsIthJotE8yodTDJkLtA) Content-type: text/html; charset=utf-8 Content-transfer-encoding: quoted-printable
I posted this as a question here:
http://stackoverflow.com/questions/6998= 713/scanning-for-unicode-numbers-in-a-string-with-d

Summarized:
The Oniguruma docs[1] seem to say that \d is = supposed to match the Unicode "Decimal_Number" category. However, in Ruby = 1.9.1 and 1.9.2 it only matches Latin 0-9 characters. Is this the correct = behavior for Ruby? Are the Oniguruma docs wrong? Am I misreading them? Do = they not apply to how Oniguruma is used within Ruby?

<= div>Test program:

#encoding: utf-8
require 'open-uri'
html =3D open("http://www.fileformat.info/in= fo/unicode/category/Nd/list.htm").read
digits =3D html.scan(/U\+= ([\da-f]{4})/i).flatten.map{ |s| s.to_i(16) }.pack('U*')

puts digits.encoding, digits
#=3D> UTF-8
#=3D= > 0123456789=D9=A0=D9=A1=D9=A2=D9=A3=D9=A4=D9=A5=D9=A6=D9=A7=D9=A8=D9=A9= =DB=B0=DB=B1=DB=B2=DB=B3=DB=B4=DB=B5=DB=B6=DB=B7=DB=B8=DB=B9=DF=80=DF=81=DF= =82=DF=83=DF=84=DF=85=DF=86=DF=87=DF=88=DF=89=E0=A5=A6=E0=A5=A7=E0=A5=A8=E0= =A5=A9=E0=A5=AA=E0=A5=AB=E0=A5=AC=E0=A5=AD=E0=A5=AE=E0=A5=AF=E0=A7=A6=E0=A7= =A7=E0=A7=A8=E0=A7=A9=E0=A7=AA=E0=A7=AB=E0=A7=AC=E0=A7=AD=E0=A7=AE=E0=A7=AF= =E0=A9=A6=E0=A9=A7=E0=A9=A8=E2=80=A6

p RUBY_DESCR= IPTION, digits.scan(/\d/)
#=3D> "ruby 1.9.2p180 (2011-02-18) = [i386-mingw32]"
#=3D> ["0", "1", "2", "3", "4", "5", "6", "7"= , "8", "9"]

Feel free to discuss here, or a= nswer on Stack Overflow if you have a solid answer and want the rep :)


[1] http://www.geocities.jp/kosa= ko3/oniguruma/doc/RE.txt

--=0A(-, /\ \/ / /\/
<= /div>= --Boundary_(ID_59lsIthJotE8yodTDJkLtA)-- --Boundary_(ID_gHyJdGyCH+UMzqa1eNhBXQ)--