From: brabuhr@... Date: 2010-11-18T14:04:10+09:00 Subject: Re: Ruby 1.9.2 UTF-8 Encoding issues whiles reading/writing files On Wed, Nov 17, 2010 at 10:31 PM, Atoli Atoli wrote: > Basically I'm trying to replace the broken sequences with "_", but the > encode method doesn't seem to do any replacements, maybe because the > forced encoding is already set to UTF-8? For the case of fixing broken files, I would probably use iconv from the shell: $ cat broken.txt 字?. $ iconv -f UTF8 -t UTF8 --byte-subst=_ broken.txt 字__. (I don't know if Ruby's Iconv module supports the subst options.) For short strings, this seems to work: irb(main):001:0> s = "\xE5\xAD\x97\xE6\x99\x2E" => "字\xE6\x99." irb(main):002:0> s.encoding => # irb(main):003:0> s.valid_encoding? => false irb(main):004:0> t = s.chars.map{|c| c.valid_encoding? ? c : '_'}.join => "字__." irb(main):005:0> t.valid_encoding? => true irb(main):006:0> t.encoding => #