From: gwtmp01@... Date: 2006-06-29T05:50:29+09:00 Subject: Re: Unicode roadmap? On Jun 28, 2006, at 4:25 PM, Jim Weirich wrote: > raise "Not PNG." unless (SINGLE_BYTE_ENCODINGS.include? > (@top.encoding) > && @top[0, 8] == "\x89PNG\x0d\x0a\x1a\x0a") > > Of course, if you can guarantee that @top is indeed a single byte > encoded BEFORE hitting this line, then the encoding test is not needed > (and I think you assume that). Does it even make sense to talk about 'encodings' in the context of binary data? I suppose you could extend the concept of encoding to capture some sort of mime type characterization of the data but isn't that a bit beyond what this thread has been talking about? I like Austin's idea of an encoding as a 'lens' with respect to the raw data. It is getting pretty hard to follow this entire discussion in the absence of some concrete examples of the imagined APIs as well as some sort of taxonomy of use cases with which to evaluate the APIs. For example: - create a copy of a text file when the text encoding is unknown - transmit a copy of a text file across a TCP/IP connection when the encoding is unknown - analyze binary data and guess at its text encoding - convert PNG image to a GIF image, in memory, to/from disk - input n characters from the keyboard/stdin/tty - count the number of words, lines, and characters in a file with an explicit encoding with an implicit encoding associated with a given locale with an implicit encoding associated with the process/thread and so on. Gary Wright