From: Yukihiro Matsumoto Date: 2008-12-15T23:51:15+09:00 Subject: [ruby-core:20574] Re: 1.9 character encoding (was: encoding of symbols) Hi, In message "Re: [ruby-core:20567] Re: 1.9 character encoding (was: encoding of symbols)" on Mon, 15 Dec 2008 20:32:48 +0900, Brian Candler writes: |Sorry, that wasn't a good example. Using #read instead: | |irb(main):001:0> File.open("/bin/sh").read(16).encoding |=> # |irb(main):002:0> File.open("/etc/passwd").read(16).encoding |=> # |irb(main):003:0> File.open("/bin/sh").gets.encoding |=> # |irb(main):004:0> File.open("/etc/passwd").gets.encoding |=> # | |So I guess: the encoding declared in open is used for #gets, but #read |always returns ASCII-8BIT. Yes, read specifies data length in bytes, so its return value should be binary (ASCII-8BIT). |I see also that inline literals have variable encoding: | |irb(main):029:0> "\x61\x62\x63".encoding |=> # |irb(main):028:0> "\x61\xc3\x9f".encoding |=> # This is old behavior. Now string literals are always in their source encoding. Try newer version. matz.