From: David Flanagan Date: 2007-11-20T09:32:34+09:00 Subject: Re: Still exists: Problems with \M-x in utf-8 encoded strings Wolfgang N��dasi-Donner wrote: > > I think the best way to solve this is not to allow "\M-." > constructs, if the resulting String object doesn't have an > encoding "", "", or > "Binary", which isn't in the list of encodings in the actual > snapshot. > > I hope I will bring better news tomorrow ;-) > > Wolfgang N��dasi-Donner > The \M escape is a weird one, but it is really uncommon. I'd be more worried about the \xxx octal escape an the \xhh hex escape, both of which allow the insertion of arbitrary bytes into a UTF-8 string. Right now when running with -Ku, any string with a byte that has the high-bit set is given UTF-8 encoding. Perhaps it would be better to make those strings binary, and to raise an error if a string literal mixes multi-byte characters with single-byte characters with the high bit set. It is possible that I introduced this behavior with my various patches for \u escapes. I thought I was just following the existing behvavior, however. Nobu and Matz, what do you think the right thing to do is here? Let me know if you want me to try to change the current behavior. David