From: Bob Hutchison Date: 2012-10-31T21:22:56+09:00 Subject: Re: REXML & HTMLentities incorrectly map to UTF-8 On 2012-10-31, at 8:19 AM, Bob Hutchison wrote: > > On 2012-10-30, at 4:30 PM, Mark S. wrote: > >> I have some XML data (UTF 8) that I'm trying to convert into another XML >> set which will be eventually UTF 16. The data contains encoded html/xml >> entities. >> >> The problem is that when I try to parse the data, html/xml entities >> inside of CDATA text are converted into 2-byte codes that don't match >> their original usage. >> >> For instance, ’: (should be right single quote) is translated into >> bytes C292 when parsed and exported and examined in a hex editor. >> >> Apparently what REXML and HTMLentities do is transliterate a value like >> "’" to character point U-146 on the Unicode chart. Unfortunately, >> this point is a CONTROL code, not a punctuation code. The real character >> point should be U-2019. >> >> Is there a fix for this? Or does one have to write their own parser to >> map these values back to appropriate usage? > > Are you saying that REXML is parsing the content of the CDATA section and replacing those entities? Or are you extracting the CDATA sections after REXML is finished and then parsing them yourself? > > If REXML is doing this then have you tried Nokogiri? (REXML should not be parsing the contents of a CDATA section) If not, "If not" --> if REXML is not parsing the entities in the CDATA section then... > then you'll need to do something along the lines of what Carlos suggested in his response. If you're still having problems can you post some sample XML and maybe some of your translation code? > > Cheers, > Bob > > >> >> Thank you, >> Mark >> >> -- >> Posted via http://www.ruby-forum.com/. >> >