From: Stefan Lang Date: 2009-02-16T20:45:53+09:00 Subject: Re: invalid byte sequence in US-ASCII (ArgumentError) 2009/2/16 Brian Candler : > Brian Candler wrote: >> Yes. IMO this is a horrendous misfeature of ruby 1.9: it asserts that >> all external data is text, unless explicitly told otherwise. Ruby must choose between treating all external data as text unless told otherwise or treat everything as binary unless told otherwise, because there is no general way to know if a file is binary or text. Given that Ruby is mostly used to work with text, it's a sensible decision to use text mode by default. Also, if you open a file with the "b" flag, it sets the files encoding to binary. You should use that flag in 1.8, too, otherwise Windows will do line ending conversion, corrupting your binary data. > And worse: the encoding chosen comes from the environment. So your > program which you developed on one system and runs correctly there may > fail totally on another. It has to default to some encoding. Your OS installation has a default encoding. It's a sane decision to use that, because otherwise many scripts wouldn't work by default on your machine. > I'm not saying that Ruby shouldn't handling encodings and conversions; > I'm just saying you should ask for them. For example: > > File.open("....", :encoding => "UTF-8") # Use this encoding Well, you can do exactly that... > File.open("....", :encoding => "ENV") # Follow the environment This is the default. > File.open("....") # No idea, treat as binary Use "b" flag, which you should do on 1.8 anyway. > I'm not going to use 1.9 without wrapper scripts to invoke Ruby with > appropriate flags to force the external encoding to a fixed value. And > that's a pain. You can set it with Encoding.default_external= at the top of your script. Stefan