From: Robert Klemme Date: 2011-03-22T00:00:28+09:00 Subject: Re: A question about Ruby 1.9's "external encoding" On Sun, Mar 20, 2011 at 6:39 PM, Brian Candler wrote: > Robert K. wrote in post #988429: >> On 20.03.2011 14:19, Brian Candler wrote: >>>> I'd say it means that the default encoding is used. >>> >>> No, it doesn't. >> >> So, which encoding is used then? > > None. Even if no encoding is used explicitly an encoding must be used nevertheless (see below). >> An encoding *has* to be used because >> you cannot write to a file without a particular encoding. > > Untrue. In Unix, read() and write() just work on sequences of bytes, and > have no concept of encoding. > > Perhaps you are thinking of a language like Python 3, where there is a > distinction between "characters" and "bytes representing those > characters" (maybe Java has that distinction too, I don't know enough > about Java to say) > > In ruby 1.9, every String is a bunch of bytes plus an encoding tag. When > you write this out to a file,  and the external encoding is nil, then > just the bytes are written, and the encoding is ignored. Which basically means that the string's own encoding is used. If you have a number of bytes and want to interpret them as characters you must use an encoding, even if it is 8 bit ASCII and there is no conversion going on. There is no such thing as a text file without encoding whether applied explicitly or not. On one side there are bytes and on the other side there are character codes (or Unicode code points). >> I could see in the console that the file was read properly. > > What you see in the console in irb does not necessarily mean much in > ruby 1.9, because STDOUT.external_encoding is nil by default too. > >> irb(main):001:0> File.open("x","w"){|io| p io.external_encoding; io.puts >> "aä"} >> nil >> => nil >> irb(main):002:0> s = File.open("x","r:UTF-8"){|io| p >> io.external_encoding; io.read} >> # >> => "aä\n" >> irb(main):003:0> s.valid_encoding? >> => true > > Now, that's more complex, and *does* show that the data is valid UTF-8. > (I wasn't arguing that it wasn't; I was arguing that your logic was > flawed, because even if the data were not valid UTF-8, your program > would have run without raising an error. Therefore the fact that it runs > without error is insufficient to show that the data is valid UTF-8) So what we learn here is that since my original string had encoding UTF-8 the encoding of the file happened to be UTF-8 as well. That basically means that by accident we can get a file with mixed encoding content. Shudder. Here's the test: >> s = "aä" => "aä" >> s.encoding => # >> s = s.encode 'ISO-8859-1' => "a\xE4" >> s.encoding => # >> Encoding.default_external => # >> $stdout.external_encoding => nil >> File.open("x","w"){|io| p io.external_encoding; io.puts(s)} nil => nil >> t = File.open("x","r:UTF-8"){|io| p io.external_encoding; io.read} # => "a\xE4\n" >> t.encoding => # >> t.valid_encoding? => false >> t.length => 3 Now let's fix it >> t.force_encoding 'ISO-8859-1' => "a\xE4\n" >> t.encoding => # >> t.valid_encoding? => true Output: >> $stdout.external_encoding => nil >> $stdout.puts t a▒ => nil >> $stdout.set_encoding($stdin.external_encoding) => #> >> $stdout.external_encoding => # >> $stdout.puts t aä => nil For me this boils down to these rules: 1. Strings are sequences of bytes 2. Strings have an associated encoding which does not need to match the actual encoding of the binary content 3. In absence of a target (external or internal, depending on direction) encoding IO operations use a String's binary data as is, otherwise they try to convert between encodings and raise an error if that is not possible. Cheers robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/