From: brabuhr@... Date: 2011-03-30T21:39:59+09:00 Subject: Re: Encoding issues when parsing HTML in 1.9 On Wed, Mar 30, 2011 at 8:25 AM, ctdev wrote: > I also tried the following on a test string: > >  s.encode("UTF-8", :invalid => :replace, :undef =>:replace, :replace > => "?") > > But it doesn't seem to replace the invalid character(s) Could that be an optimization in encode: since the string is already thought to be UTF-8, just return it? > s = "PROT\xC9GÉ"=> "PROT\xC9G\u00C9" > s.encode("UTF-8", :invalid => :replace, :undef =>:replace, :replace => "?") => "PROT\xC9G\u00C9" > s. encode('ISO8859-9', :invalid => :replace, :undef =>:replace, :replace => "#"). encode("UTF-8", :invalid => :replace, :undef =>:replace, :replace => "?") => "PROT#G\u00C9"