From: Matthew Kerwin Date: 2012-11-05T22:09:09+09:00 Subject: Re: REXML & HTMLentities incorrectly map to UTF-8 --f46d042ef689d3ec7f04cdbf30a8 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable Something that might not have been noticed: U+0092 (apparently called "PRIVATE USE TWO" but usually a rounded apostrophe) when encoded in UTF-8 is two bytes long: 0xC2 0x92. Where you're seeing what appears to be U+C292, I would assume you're actually seeing a two-byte UTF-8 encoded form of U+92 (remember that UTF-8 is emphatically not UCS-2). Thus the character would be being interpreted "correctly" as the apostrophe char, and output as UTF-8. If it starts looking like "=C3=82=E2=80=99" (that is, 0xC3 0x82 0xC2 0x92) = then you're in double-encoding land. And for the record, U+C292 isn't a control code, it's a hangul character = =EC=8A=92 On 5 November 2012 22:53, Kouhei Sutou wrote: > Hi, > > In > "Re: REXML & HTMLentities incorrectly map to UTF-8" on Mon, 5 Nov 2012 > 11:18:33 +0900, > "Mark S." wrote: > > > Kouhei Sutou wrote in post #1082578: > >> Could you show me a sample Ruby code? > >> If I can reproduce your problem with the code on my machine, I > >> will fix the problem and the fix will be shipped in Ruby 2.0.0. > > > > Here is some code to produce the problem, plus the input xml and the > > output xml that I got when running the code. If you view the output in > > an editor that shows hex code, you'll see that the apostrophe in > > "fund's" becomes transliterated to char point C292 -- which is just an > > unused control code. > > > > The entity code used for the apostrophe is ’ which my Oreilly HTML > > book indicates should indeed be rendered as an apostrophe. > > Thanks for providing sample code. > > First, "’" should be handled as U+0092 in XML. > See also: > http://www.w3.org/TR/REC-xml/#sec-references > > If the character reference begins with " &#x ", the digits > and letters up to the terminating ; provide a hexadecimal > representation of the character's code point in ISO/IEC > 10646. If it begins just with " &# ", the digits up to the > terminating ; provide a decimal representation of the > character's code point. > > In your case, "&#" case. It means that 146 is handled as > decimal and it is 0x92 in hexadecimal. So ’ is U+0092 > in XML. > > (Note that XML is not HTML.) > > > But the problem is even worse. It turns out that if there is any HTML > > tagging inside of the CDATA ... REXML deletes the data! Sometimes it > > even hangs up with a "tree parsing error" (not exact text) with no > > indication what source tag is giving the problem. (Sorry, can't provide > > that sample input since its 15megs of semi-private data). > > I can't reproduce your problem with the following script: > > require "rexml/document" > > document =3D REXML::Document.new(<<-EOX) > > tag]]> > > EOX > > note =3D document.elements["/notebook/note"] > cdata =3D note[0] > p cdata > # =3D> "tag" > > It seems that the output includes HTML tag in CDATA. > > > Thanks, > -- > kou > > --=20 Matthew Kerwin, B.Sc (CompSci) (Hons) http://matthew.kerwin.net.au/ ABN: 59-013-727-651 "You'll never find a programming language that frees you from the burden of clarifying your ideas." - xkcd --f46d042ef689d3ec7f04cdbf30a8 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable

Something that might not have been noticed:

U+0092 (apparently called "PRIVATE USE TWO" but usually a roun= ded apostrophe) when encoded in UTF-8 is two bytes long: 0xC2 0x92. =C2=A0W= here you're seeing what appears to be U+C292, I would assume you're= actually seeing a two-byte UTF-8 encoded form of U+92 (remember that UTF-8= is emphatically not UCS-2). =C2=A0Thus the character would be being interp= reted "correctly" as the apostrophe char, and output as UTF-8.

If it start= s looking like "=C3=82=E2=80=99" (that is,=C2=A00xC3 0x82 0xC2 0x= 92) then you're in double-encoding land.

And for the record, U+C292 isn't a control code, it's a hangul char= acter=C2=A0=EC=8A=92

On 5 November 2012 2= 2:53, Kouhei Sutou <kou@cozmixng.org> wrote:
Hi,

In <f= ec82c6b596b842bf6731e87991ed3cc@ruby-forum.com>
=C2=A0 "Re: REXML & HTMLentities incorrectly map to UTF-8" on= Mon, 5 Nov 2012 11:18:33 +0900,
=C2=A0 "Mark S." <lists@ruby-forum.com> wrote:

> Kouhei Sutou wrote in post #1082578:
>> Could you show me a sample Ruby code?
>> If I can reproduce your problem with the code on my machine, I
>> will fix the problem and the fix will be shipped in Ruby 2.0.0. >
> Here is some code to produce the problem, plus the input xml and the > output xml that I got when running the code. If you view the output in=
> an editor that shows hex code, you'll see that the apostrophe in > "fund's" becomes transliterated to char point C292 -- wh= ich is just an
> unused control code.
>
> The entity code used for the apostrophe is &#146; which my Oreilly= HTML
> book indicates should indeed be rendered as an apostrophe.

Thanks for providing sample code.

First, "&#146;" should be handled as U+0092 in XML.
See also:
=C2=A0 http://www.w3.org/TR/REC-xml/#sec-references

=C2=A0 If the character reference begins with " &#x ", the di= gits
=C2=A0 and letters up to the terminating ; provide a hexadecimal
=C2=A0 representation of the character's code point in ISO/IEC
=C2=A0 10646. If it begins just with " &# ", the digits up to= the
=C2=A0 terminating ; provide a decimal representation of the
=C2=A0 character's code point.

In your case, "&#" case. It means that 146 is handled as
decimal and it is 0x92 in hexadecimal. So &#146; is U+0092
in XML.

(Note that XML is not HTML.)

> But the problem is even worse. It turns out that if there is any HTML<= br> > tagging inside of the CDATA ... REXML deletes the data! Sometimes it > even hangs up with a "tree parsing error" (not exact text) w= ith no
> indication what source tag is giving the problem. (Sorry, can't pr= ovide
> that sample input since its 15megs of semi-private data).

I can't reproduce your problem with the following script:

=C2=A0 require "rexml/document"

=C2=A0 document =3D REXML::Document.new(<<-EOX)
=C2=A0 <notebook>
=C2=A0 =C2=A0 <note><![CDATA[<html>tag</html>]]><= ;/note>
=C2=A0 </notebook>
=C2=A0 EOX

=C2=A0 note =3D document.elements["/notebook/note"]
=C2=A0 cdata =3D note[0]
=C2=A0 p cdata
=C2=A0 # =3D> "<html>tag</html>"

It seems that the output includes HTML tag in CDATA.


Thanks,
--
kou




--
=C2=A0 Matth= ew Kerwin, B.Sc (CompSci) (Hons)
=C2=A0 http://matthew.kerwin.net.au/
=C2=A0 ABN= : 59-013-727-651

=C2=A0 "You'll never find a programming language that frees=C2=A0 you from the burden of clarifying your ideas." - xkcd
--f46d042ef689d3ec7f04cdbf30a8--