From: Ian Macdonald Date: 2007-02-14T08:43:06+09:00 Subject: Re: Puzzling regex behaviour On Wed 14 Feb 2007 at 08:01:15 +0900, David Balmain wrote: > On 2/14/07, David Balmain wrote: > >On 2/14/07, Ian Macdonald wrote: > >> However, I'm still puzzled why there's a difference between irb and a > >> stand-alone script. > > > >Maybe your editor saves the script in UTF-8 format. The irb example > >clearly encodes the string in ISO-8859-1. That could explain the > >difference. > > For example; > > ~$ echo $LANG > en_US.ISO-8859-1 > ~$ irb -f > irb(main):001:0> "pr\351f\351r\351es" =~ /[^[:alnum:]]/ > => nil > irb(main):002:0> "pr\303\251f\303\251r\303\251es" =~ /[^[:alnum:]]/ > => 3 > > Not exactly what you had but it probably has something to do with the > encoding of the �. My editor is vim and I run it in the nl_NL locale, so it doesn't start in UTF-8 mode. To double-check: :set encoding? encoding=latin1 And if we dump my little script: $ od -c foo 0000000 # ! / u s r / b i n / r u b y 0000020 - w \n \n f o o = " p r 351 f 351 0000040 r 351 e s " \n p f o o = ~ / 0000060 [ ^ [ : a l n u m : ] ] / \n p 0000100 f o o = ~ / \ W / \n You can see that it is, indeed, saved as Latin-1, not UTF-8. The mystery continues. ;-) Ian -- Ian Macdonald | It's not whether you win or lose, it's how ian@caliban.org | you place the blame. http://www.caliban.org/ | | |