From: Austin Ziegler Date: 2006-06-26T15:20:08+09:00 Subject: Re: Unicode roadmap? On 6/26/06, Izidor Jerebic wrote: > On 26.6.2006, at 5:01, Yukihiro Matsumoto wrote: > > I still don't see how separate types and behaviors would be more > > logical and break far less. For example, if I want to check EXIF > > conformance of a jpeg file, I do > > > > def self.exif_file? (filename) > > exif_header = "\xff\xd8\xff\xe1" > > magic = File.open(filename) {|f| f.read(4) } > > magic == exif_header > > end > > > > I am not sure what you expect about separation, but I doubt separation > > would make above code to "be more logical and break far less". > Above code assumes all file operations return byte arrays. What is > the code when we want to obtain String of characters? As Tim Bray pointed out in a response to me, trying to get a String from a file is a ludicrous operation. I was mocking the API required (e.g., File#read_string or something equally bozonic). You need to read your data and *then* mark it as a String with a particular encoding. And if you *globally* change the interpretation of File#read to be String, you will be breaking the ability to read truly binary data. > The problem is the auto-magic encoding handling which is required to > have text processing be as simple as it is now. You can have either > text processing (which adds encoding handling for us, combines bytes > in characters etc.) or byte processing (which does not). How do we > distinguish between the two modes of operation? > > The obvious way is by adding a ByteArray. But maybe there is better > way... Yes. It's to actually read what has been suggested. The m17n String won't be a magic bullet. But you'll be able to do something like: bv = File.open("file.txt", "rb") { |f| f.read } sv = bv.with_encoding(:utf8) Or something like that. And you can still do bv == "\xff\xd8\xff\xe1" as appropriate. -austin -- Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ * austin@halostatue.ca * http://www.halostatue.ca/feed/ * austin@zieglers.ca