From: Brian Candler Date: 2008-12-15T23:50:11+09:00 Subject: Re: ruby 1.9.1: Encoding trouble: broken US-ASCII String James Gray wrote: >> I would have thought that a US-ASCII regexp should be able to match >> ISO-8859-1 data, and perhaps vice versa, but it seems not. > > It does: > > $ ruby_dev -e 'p "r�sum�".encode("ISO-8859-1") =~ /foo/' > nil > $ ruby_dev -e 'p "r�sum� foo".encode("ISO-8859-1") =~ /foo/' > 7 I found that too, but was confused by the "broken US-ASCII string" exception which the OP saw. I suppose the external_encoding is defaulting to US-ASCII on that system. This means his program will break on every file passed into it which has a character with the top bit set. You can argue that's "failsafe", in the sense of bombing out rather than continuing processing with the wrong encoding, and it therefore forces you to change your program or the command-line args to specify the actual encoding in use. However, that's pretty unforgiving. I can use Unix grep on a file with unknown character set or broken UTF-8 characters and it works quite happily. Wouldn't it be kinder to default to BINARY if the encoding is unspecified? irb(main):011:0> s = "foo\xff\xff\xffbar".force_encoding("BINARY") => "foo\xFF\xFF\xFFbar" irb(main):012:0> s =~ /foo/ => 0 >> Maybe what's really needed is a sort of "anti-/u" option which means >> "my >> regexp literals are meant to match byte-at-a-time, not >> character-at-a-time" > > That's what BINARY means. On the String side, yes. I was thinking of an option on the Regexp: /foo/b or somesuch. (In contrast to /foo/u in 1.8 meaning 'this Regexp matches unicode') Or you can you set BINARY encoding on the Regexp too? I couldn't see how. -- Posted via http://www.ruby-forum.com/.