From: James Gray Date: 2008-09-18T03:35:53+09:00 Subject: Re: How to clean an xml files from non-utf-8 chars? On Sep 17, 2008, at 11:47 AM, Jeremy Hinegardner wrote: > On Wed, Sep 17, 2008 at 09:44:23PM +0900, James Gray wrote: >> On Sep 17, 2008, at 4:07 AM, Krzysieq wrote: >> >>> I have a problem. I'm trying to parse with ruby some test results >>> from >>> jmeter, that are stored in xml files. Unfortunately, while they >>> should be >>> utf-8, some of them aren't. Probably because some db data isn't. >>> In any >>> case, this makes other toys break down, like xslt transformation and >>> anything else that relies on the xml files being utf-8. >>> >>> Does anyone know, how to get rid of such characters? >> >> If you can figure out the encoding they are actually in, I >> recommend using >> Iconv's transliterate mode: >> >> require "iconv" >> Iconv.conv("UTF-8//TRANSLIT", old_encoding_name, data) > > This is the approach we have take on some of our code, basically we > wanted to > replicate the 'iconv -c' behavior. Does TRANSLIT do this ? I've > never used > that mode before. //TRANSLIT is better than that. It tries to translate the characters. Thus a UTF-8 ellipse would become three periods if converted to ISO-8859-1 with //TRANSLIT. You can mimic -c though, just use //IGNORE instead of //TRANSLIT. You can even do //TRANSLIT//IGNORE which transliterates what it can and discards the rest. James Edward Gray II