From: Austin Ziegler Date: 2006-06-27T06:52:14+09:00 Subject: Re: Unicode roadmap? On 6/26/06, Jim Weirich wrote: > Thanks for the response, Austin. It seemed to help clearify the issues > (at least for me). > > Austin Ziegler wrote: > > d1 = File.open("file.txt", "rb", encoding: :utf8) { |f| f.read } > Question: Does the encoding parameter specify the encoding of the file, > or the encoding of the strings you get back (my guess is both). I would assume both, based on what I've seen from Matz. > Related question: In environments that use a lot of different encodings, > are there ways or conventions for specifying the encoding, or do you > just have to "know". In my experience, you just have to "know" unless you can do some detection of the encoding. I think that only UTF-16 or UTF-32 is really amenable to this ;) This is one of the problems that I've seen with the encoding work that I've done. If I'm reading a list of files from a NetWare server, what encoding is the data in? I don't necessarily have a Unicode interface -- and my code page may not match the server's code page. *Whenever* you're dealing with legacy data, you have to "agree" or guess and hope you're right. >> s1.encoding = :utf8 > Another Question: When you set the encoding, are you: > > (A) Just changing the encoding specifier without changing the > underlaying string. > (B) Re-encoding the string according to the new encoding specifier. > (B) seems to be implied by the attribute notation, but that seems a bit > dangerous in my mind. I personally consider it to be (A) because I believe that encoding is a lens. If you want (B) it should be s1.recode(:utf8). But #recode would not work on an encoding of "binary" (or "raw"); #recode would be similar to the Iconv steps you would use today. -austin -- Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ * austin@halostatue.ca * http://www.halostatue.ca/feed/ * austin@zieglers.ca