From: ts Date: 2006-03-24T01:47:32+09:00 Subject: Re: perl regexp to ruby one conversion ? >>>>> "U" == =?ISO-8859-1?Q?Une b=E9vue?= writes: U> the above regexp doesn't work as expected with ruby, i've compared the U> output for the same files with perl and ruby, ruby says always "yes it U> is UTF-8", where perl says NO over an ISO-8859-1 encoded file... (even U> after wipping out the first line the first ^and the last $) moulon% cat b.rb field='&���!����tybvn�' utf8rgx=Regexp.new('^( [\x09\x0A\x0D\x20-\x7E] # ASCII | [\xC2-\xDF][\x80-\xBF] # non-overlong 2-byte | \xE0[\xA0-\xBF][\x80-\xBF] # excluding overlongs | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2} # straight 3-byte | \xED[\x80-\x9F][\x80-\xBF] # excluding surrogates | \xF0[\x90-\xBF][\x80-\xBF]{2} # planes 1-3 | [\xF1-\xF3][\x80-\xBF]{3} # planes 4-15 | \xF4[\x80-\x8F][\x80-\xBF]{2} # plane 16 )*$', Regexp::EXTENDED) p utf8rgx =~ field moulon% moulon% file b.rb b.rb: ISO-8859 text moulon% moulon% ruby b.rb nil moulon% Guy Decoux