From: Dominik Bathon Date: 2006-03-24T06:37:30+09:00 Subject: Re: perl regexp to ruby one conversion ? Hi, On Thu, 23 Mar 2006 19:13:51 +0100, "Une b�vue" wrote: > utf8rgx=Regexp.new('^( > [\x09\x0A\x0D\x20-\x7E] # ASCII > | [\xC2-\xDF][\x80-\xBF] # non-overlong 2-byte > | \xE0[\xA0-\xBF][\x80-\xBF] # excluding overlongs > | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2} # straight 3-byte > | \xED[\x80-\x9F][\x80-\xBF] # excluding surrogates > | \xF0[\x90-\xBF][\x80-\xBF]{2} # planes 1-3 > | [\xF1-\xF3][\x80-\xBF]{3} # planes 4-15 > | \xF4[\x80-\x8F][\x80-\xBF]{2} # plane 16 > )*$', Regexp::EXTENDED) As I understand it utf8rgx matches any string that is utf8, which includes pure ascii strings (see first line). So it should match http://www.rubycentral.com/book/tut_exceptions.html. First, here is a working version: $ cat utf8tst.rb utf8rgx = /\A( [\x09\x0A\x0D\x20-\x7E] # ASCII | [\xC2-\xDF][\x80-\xBF] # non-overlong 2-byte | \xE0[\xA0-\xBF][\x80-\xBF] # excluding overlongs | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2} # straight 3-byte | \xED[\x80-\x9F][\x80-\xBF] # excluding surrogates | \xF0[\x90-\xBF][\x80-\xBF]{2} # planes 1-3 | [\xF1-\xF3][\x80-\xBF]{3} # planes 4-15 | \xF4[\x80-\x8F][\x80-\xBF]{2} # plane 16 )*\z/x p utf8rgx === ARGF.read $ curl -s http://www.linux-france.org/ | ruby utf8tst.rb false $ curl -s http://www.rubycentral.com/book/tut_exceptions.html | ruby utf8tst.rb true Your problem was that in Perl ^ and $ only match beginning and end of string, but in ruby they also match beginning and end of line. So if a string contains for example a single empty line, it does always match: irb(main):001:0> a = "xxx\n\nyyyy" => "xxx\n\nyyyy" irb(main):002:0> a =~ /^(w)*$/ => 4 So for beginning and end of string in ruby you need \A and \z: irb(main):003:0> a =~ /\A(w)*\z/ => nil Hope that helps, Dominik