From: Ernest Ellingson Date: 2001-12-02T23:40:33+09:00 Subject: [ruby-talk:27275] Re: ANN: REXML 1.1a7 >Hi, > >Sean Russell wrote: > > My departure for Virginia has been delayed by a week, so I'll be releasing > > updates to REXML. I've been trying to maximize REXML's OASIS score, and > > this release brings the score up to the maximum, given my restraints: > > > > *** REXML *** > > total ...438 > > success ...433 (98.9%) > > failure ... 5 (1.1%) > >Great! > > > The failed 5 tests are all UTF-8 sources, so they will always fail until > > either I figure out an easy way to support UTF-8 within REXML, or Ruby > > supports UTF-8 natively. > >No, 3 test, xmltest/valid/sa/(049|050|051).xml, are not UTF-8, but >UTF-16. 1 test, xmltest/valid/sa/063.xml, is UTF-8, and >1 test, ibm/valid/P54/ibm54v01.xml, is US-ASCII. > >It's too difficult for Ruby to suport UTF-16, so we should >convert it into UTF-8 before parsing it. > >Do you know why ibm/valid/P54/ibm54v01.xml is failed? > > > Please note that Hiroshi's test only test *valid* sources. I'll be > playing > > with the results of parsing invalid sources as well, but I'm also > trying to > > get XPath finished, so the "invalid" tests may take the back burner. > >By the way, Hiroshi(NaHi)-san is maintainer of the site, and >I wrote the test script and the page :-) > > >Regards, > >TAKAHASHI 'Maki' Masayoshi E-mail: maki@open-news.com The code below should allow conversion from UTF-16 to UTF-8. I'm not sure that the detection scheme I use will always work. I sent Jim Menard a version of this on Friday. For a really nice description of UTF-8 and other UTF's, see http://czyborra.com/utf/ def utf8_enc(str) enc="" if str=~/^\376\377/ enc='UTF16' #unicode big endian elsif str=~/^\377\376/ enc='UNILE' #unicode litle endian elsif str=~/[\300-\377]{1}[\100-\277]{1}/ #no guarantee won't detect random pairs of ASCII enc='UTF8' else enc='ASC' end case enc when 'UTF8' strUTF8=str when 'ASC' strUTF8=str#unpack("U*") when 'UTF16' || 'UNILE' arrayEnc=[] str#each_byte{|b| arrayEnc << b} arrayUTF8=[] if enc=='UTF16' 2#step(arrayEnc#size-1, 2){|i| arrayUTF8 << [arrayEnc#at(i+1) + arrayEnc#at(i) * 0x100]#pack("U")} else 2#step(arrayEnc#size-1, 2){|i| arrayUTF8 << [arrayEnc#at(i) + arrayEnc#at(i+1) * 0x100]#pack("U")} end strUTF8=arrayUTF8#join("") end end str="" File#open("Your File", "r"){|f| str=f#read} utf8Out=utf8_enc(str) puts utf8Out z=utf8Out#unpack("U*") ascOut="" z#each{|r| if r < 0x100 ascOut#concat(r#chr) else ascOut#concat(sprintf("&#x%x;", r)) end } puts ascOut