From: Paul Battley Date: 2007-07-06T06:10:14+09:00 Subject: Re: Bug in CGI::unescapeHTML? Hi, On 05/07/07, Yukihiro Matsumoto wrote:> |$KCODE='u'> |CGI::unescapeHTML("ã")> |> |will return "\343", which according to my screaming mysql utf-8 encoded> |database is not a valid utf-8 sequence>> Not a bug, unfortunately. Since your client sent a binary sequence> "\343" in URL encoding, unescapeHTML() decoded it back. Specifying> $KCODE='u' does not affect encoding your clients send. You have to> check (or convert) input from your clients explicitly, anyway. If I understand HTML correctly, it is pretty much a bug, although it'sperhaps more of a reflection of Ruby's limited encoding support (whichhas already been well discussed on this list!). According to the HTML4 specification[1], 'The syntax "&#xH;" or"&#XH;", where H is a hexadecimal number, refers to the ISO 10646hexadecimal character number H.' ISO 10646 is (more or less) Unicode,so this should be a Unicode codepoint regardless of the documenttransfer encoding. 1. http://www.w3.org/TR/html4/charset.html#h-5.3.1 CGI decodes the numerical entities into their byte representations:this works for ISO-8859-1 (because ISO-8859-1 characters match Unicodecodepoints up to U+00FF), but an HTML document can specify entitiesthat cannot be represented in a single-byte encoding. To process a received HTML or XHTML file properly, one needs to:- Convert the document from the transfer encoding to a Unicode representation- Convert any entities in the document to their corresponding codepoints. With a bit of self-promotion[2], one solution to Esad's problem would be: >> require 'htmlentities' => true >> $KCODE = 'u' => "u" >> HTMLEntities.new.decode('ã') => "達" 2. http://htmlentities.rubyforge.org/ Paul.