From: Eric Hodel Date: 2009-07-23T07:45:23+09:00 Subject: Re: Byte–stream parsing in Ruby On Jul 22, 2009, at 00:46, Brian Candler wrote: > Haha, fooled you. You thought it was safe to append a non-UTF8 > character > to a UTF8 string (after all, you did before quite happily), but this > time you get an exception. So now you have to do: > >>> str.force_encoding("ASCII-8BIT") > => "hello\xFF" >>> str << 250.chr > => "hello\xFF\xFA" >>> str.force_encoding("UTF-8") > => "hello\xFF\xFA" >>> str.valid_encoding? > => false > > This is why I hate ruby 1.9. I don't think that's a valid UTF-8 byte sequence... > Incidentally, I needed to do something similar in ruby-1.8 recently, > and > it was very straightforward. > > def is_utf8?(str) > Iconv.iconv('UTF-8','UTF-8',str) > true > rescue Iconv::IllegalSequence > false > end Oh, I see there's another tool let's try it! $ cat conv.rb str = "\xFF\xFA" require 'iconv' converted = Iconv.iconv 'UTF-8', 'UTF-8', str puts converted $ ruby -v conv.rb ruby 1.8.6 (2008-08-11 patchlevel 287) [universal-darwin9.0] conv.rb:6:in `iconv': "\377\372" (Iconv::IllegalSequence) from conv.rb:6 Ok, so it's not valid. Let's get a valid byte sequence... $ cat conv.rb str = "\xE2\x98\x83" require 'iconv' converted = Iconv.iconv 'UTF-8', 'UTF-8', str puts converted $ ruby conv.rb ☃ Ok, so that works! Now let's use 1.9's built-in encoding stuff with our valid byte sequence: $ cat conv.rb # encoding: utf-8 str = "hello " p :encoding => str.encoding str << 0xE2.chr str << 0x98.chr str << 0x83.chr puts str $ ruby19 conv.rb {:encoding=>#} hello ☃ huh, it worked fine. So you're mad that Ruby doesn't let you shoot yourself in the foot?