From: Nathan Beyer Date: 2012-09-21T08:43:41+09:00 Subject: Re: Encoding question On Thu, Sep 20, 2012 at 2:58 AM, Brian Candler wrote: > Nathan Beyer wrote in post #1076380: >> Once you do know the encoding, the proper >> way to open up a file is to declare the external encoding of the file, >> like this - >> >> f = File.open('somefile.txt', 'r:iso-8859-1') > > That is correct. > >> Then when you read content from the file, the data in the file will be >> transcoded from 'iso-8859-1' to the default internal encoding of the >> Ruby interpreter instance (generally UTF-8). > > That is incorrect. But encodings in ruby are such a damned mess that I'm > not surprised that many people don't understand it. Indeed ... 1.9.3-p194 :015 > f = File.open('hello.txt', 'r:windows-1252') => # 1.9.3-p194 :016 > c = f.read => "hello world!" 1.9.3-p194 :017 > c.encoding => # Looking at the documentation, I see what I missed (http://ruby-doc.org/core-1.9.3/IO.html#method-c-new). I suppose that makes. > > *No* transcoding takes place in the above example, regardless of the > where it runs; the data is read into the string as-is. However, every > String which you read from the file using f.gets or f.getc is marked to > say that it is encoded using UTF-8. (But not Strings which you read > using f.read) I'm not seeing this behavior. 1.9.3-p194 :001 > f = File.open('hello.txt', 'r:windows-1252') => # 1.9.3-p194 :002 > c = f.getc => "h" 1.9.3-p194 :003 > c.encoding => #