From: "Wolfgang Nádasi-Donner" Date: 2007-10-22T21:44:16+09:00 Subject: Re: Encoding::primary_encoding Nobuyoshi Nakada schrieb: > Wolfgang N�dasi-Donner wrote in [ruby-core:12823]: >> Michal Suchanek schrieb: >>> On 22/10/2007, Wolfgang N�dasi-Donner wrote: >>>> Nobuyoshi Nakada schrieb: >>>>> In general, by the magic coding comment. >>>> I still don't understand this. Are these "magic coding comment"s set by the >>>> editors? - I don't have something like this when using SciTE on Windows (I'm >>>> usually a SciTE user). >>> No, these are comments that tell the editor what encoding it should >>> use (like BOM but more general). Some editors accept some format of >>> these comments or another. Some have scripts that can automatically >>> generate them. The Ruby parser was just modified to take advantage of >>> several flavours of these comments. >> Doesn't this mean, that Ruby has in case of my sources (SciTE) only can decide >> on basis of a BOM (which may be completely valid in binary data) or some heuristics? > Just assumes ASCII. Unfortunately my knowledge about "EUC-JP" and "Shift_JIS" is best expressed by "nil", so I don't know if there is something like a "BOM" for data identification. For Ascii, extended-Ascii (eight valid bits, e.g ISO-8859-1), and UTF-8 I understand, that the following is true for Ruby 1.9(.0) in the moment and will be true for Ruby 1.9.x (x>=1): 1 - If a "magic coding comment" is found inside the Ruby source (produced by an editor or inserted manually), the source will be expected written in the given encoding (this means, that the "magic coding comment" must have same bit pattern in all possible encodings). 2 - If there is no "magic coding comment" found in the source, but the file begins with an UTF-8 BOM (0xEF, 0xBB, 0xBF), the source will be processed as UTF-8, and the BOM will be ignored. 3 - If there is neither a "magic coding comment" nor an UTF-8 BOM in the source, it will be expected as ASCII, where characters not in range (0x00..0x7F) are invalid outside String constants and comments. Five questions: - Is this the correct sequence ("magic coding comment" before looking for an UTF-8 BOM)? - Are there "magic byte sequences" for "EUC-JP" and "Shift_JIS"? - Does a complete list of possible "magic coding comment"s exist? - Where? - Are there wrong parts inside my description? - Will this behaviour stay stable in Ruby 1.9.x (x>=1)? I would like to put this information as soon as possible together for documentation purposes. Wolfgang N�dasi-Donner