From: Robert Klemme Date: 2010-02-24T07:10:04+09:00 Subject: Re: [ENCODING] UTF8 hell On 23.02.2010 12:10, Xavier Noëlle wrote: > 2010/2/2 Robert Klemme : >> You probably first want to find out whether the byte sequence is valid >> UTF-8 or not. For that you would need to look at the bytes in the >> String. I guess chances are that your String's byte sequence is NOT >> valid UTF-8 OR you have a character in the string that has no >> lowercase representation. > I dug into the problem and ended up with this line: self.force_encoding('UTF-8') > Believing that the string #encoding was right was a wrong choice, then > I assumed the database provided valid UTF8 strings. The string you show below does not look like UTF-8 encoded, probably rather ISO-8859-1 or such. If you enforce an encoding you leave the byte sequence untouched. This leads to the kind of error you describe below. > BUT (because, there's a but...), for some reason I don't understand, > some strings are unwilling to work: > > Example: > puts self => médicals > self.each_byte {|b| print "#{b} "} => 109 233 100 105 99 97 108 115 > > 233 is, AFAIK, a valid UTF8 character, but calling gsub(anything) (eg. > self.gsub('ruby', 'zorglub')) on this string leads to: `gsub': invalid > byte sequence in UTF-8 (ArgumentError). > > Where am I wrong ? As far as I can see 233 starts a three byte sequence http://en.wikipedia.org/wiki/UTF-8#Description I did not dig deeper but it may be that by forcing UTF-8 on an ISO something encoded string you broke it. Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/