From: Gavin Kistner Date: 2011-08-10T07:21:12+09:00 Subject: Re: Is \d supposed to match Unicode Numbers? --Boundary_(ID_K8L6IIU++9pzaADbKQBerw) Content-type: text/plain; charset=ISO-8859-1; format=flowed Content-transfer-encoding: quoted-printable On Aug 09, 2011, at 03:38 PM, Brian Candler wrote:=0A= Gavin Kistner wrote in post #1015799:=0A> Is this the correct behavior for= =0A> Ruby? Are the Oniguruma docs wrong? Am I misreading them? Do they not= =0A> apply to how Oniguruma is used within Ruby?=0A=0A* \w and \d match on= ly Latin letters and digits=0A* [[:alpha:]] and [[:digit:]] match the full= unicode set=0A=A0=0ADefinitely helpful in achieving the end goal - thanks= !=0A=0AAny guess as to how to reconcile this behavior with what the Onigur= uma "ONIG_SYNTAX_RUBY" document says? Looking at the secton on \w, we may = have a clue:=0A=0A=A0 =A0 \w =A0word character=0A=A0 =A0 =A0 =A0 Not Unico= de:=A0alphanumeric, "_" and multibyte char.=A0=0A=A0 =A0 =A0 =A0 Unicode:=A0= General_Category -- (Letter|Mark|Number|Connector_Punctuation)=0A=A0 =A0 [= ..]=0A=A0 =A0 \d =A0decimal digit char=0A=A0 =A0 =A0 =A0 Unicode: General= _Category -- Decimal_Number=0A=0APerhaps "Not Unicode" means "this is how = it behaves in some non-Unicode mode", and "Unicode" means "this is how it = behaves in some Unicode mode". And perhaps missing from the doc for \d is = something like "Not Unicode: 0-9".=0A=0AIf that is correct then the next q= uestion for me is how to enable Unicode-mode for Oniguruma. The /u flag on= a regexp does not do it, since:=0A=0A=A0 =A0 "ab=E7".scan(/\w/) =3D=3D "a= b=E7".scan(/\w/u)=0A=0A=0A= --Boundary_(ID_K8L6IIU++9pzaADbKQBerw) Content-type: multipart/related; boundary="Boundary_(ID_aBewvNv+ckYGMKM2OtfyUA)"; type="text/html" --Boundary_(ID_aBewvNv+ckYGMKM2OtfyUA) Content-type: text/html; charset=ISO-8859-1 Content-transfer-encoding: quoted-printable
On Aug 09, 2011, at 03:38 PM, Brian Candler <b.candler@pobox.com&g= t; wrote:
Gavin Kistner wrote in post #1015799:
=0A> Is t= his the correct behavior for
=0A> Ruby? Are the Oniguruma docs wrong= ? Am I misreading them? Do they not
=0A> apply to how Oniguruma is u= sed within Ruby?
=0A
* \w and \d match only Latin letters and digits=
=0A* [[:alpha:]] and [[:digit:]] match the full unicode set
 
Definitely helpful in ac= hieving the end goal - thanks!

Any guess as to how to reconcile this behavior with what the Onigu= ruma "ONIG_SYNTAX_RUBY" document says? Looking at the secton on \w, we may= have a clue:

 = ;   \w  word character
       = ; Not Unicode: alphanumeric, "_" and multibyte char. 
= &= nbsp;       Unicode: General_Category -- (Letter|Mark|= Number|Connector_Punctuation)
    [...]
    \d  decimal digit char
    &nbs= p;   Unicode: General_Category -- Decimal_Number
Perhaps "Not Unicode" means "this is how it behaves in some no= n-Unicode mode", and "Unicode" means "this is how it behaves in some Unico= de mode". And perhaps missing from the doc for \d is something like "Not U= nicode: 0-9".

If that is correct then the next qu= estion for me is how to enable Unicode-mode for Oniguruma. The /u flag on = a regexp does not do it, since:

    "= ab=E7".scan(/\w/) =3D=3D "ab=E7".scan(/\w/u)


=
= --Boundary_(ID_aBewvNv+ckYGMKM2OtfyUA)-- --Boundary_(ID_K8L6IIU++9pzaADbKQBerw)--