From: Michael Fellinger Date: 2010-02-24T13:12:56+09:00 Subject: Re: [ENCODING] UTF8 hell On Wed, Feb 24, 2010 at 12:18 AM, Xavier Noëlle wrote: > 2010/2/23 Yukihiro Matsumoto : >> 233 is not a valid UTF-8 character.  The byte sequence for médicals is >> <109 195 169 100 105 99 97 108 115>. > > Indeed. In the meantime, I changed the code with this one: > def isUTF8() >  begin >    self.unpack('U*') >  rescue >    return false >  end >  return true > end > > if isUTF8() >  self.force_encoding('UTF-8') > else >  self.force_encoding('ISO-8859-1') >  self.encode!('UTF-8') > end string = "\xE8te pour luth" # "\xE8te pour luth" string.encoding # # string.valid_encoding? # false string.force_encoding('ISO-8859-1') # "ète pour luth" string.valid_encoding? # true string.upcase # "èTE POUR LUTH" > This (ugly) quickfix works for what I need, but I don't know if this > problem can be somehow resolved in another way. The problem being that > my SQL database has a VARBINARY column with an unknown encoding. Is > there a way to deal with the various possible encoding or to ask MySQL > to return UTF8 converted data, or is it necessary to clean data before > inserting them ? > > -- > Xavier NOELLE > > -- Michael Fellinger CTO, The Rubyists, LLC 972-996-5199