From: Nikolai Weibull Date: 2005-06-16T07:36:39+09:00 Subject: Re: Regex help Ezra Zygmuntowicz wrote: > Could someone help me do a little regex conversion? I've got a > few perl compatible regexes from a php script I am trying to port to > ruby but I need a little help. Here are the php functions: > > $buffer = preg_replace("#(? [a-zA-Z]{2,4}(?:/[^ \n\r\"\'<]+)?#", "http://$0", $buffer); > $buffer = preg_replace("#(? (?:http:\/\/|https:\/\/|ftp:\/\/)(?:[a-zA-Z0-9\-]+\.)+[a-zA-Z]{2,4} > (?::[0-9]+)?(?:/[^ \n\r\"\'<]+)?#", " > $0", $buffer); > $buffer = preg_replace("#(?<=[\n ])([a-z0-9\-_.]+?)@([^,< \n\r]+)#i", > "$0", $buffer); OK, this wins my newly instated prize for _worst regexes ever_. Inefficient, inconclusive, inconsistent, and just plain wrong. I really hope you don’t have to work with a lot of code like this. Nonetheless, here’s my solution: domain = /(?:[[:alnum:]\-]+\.)/ tld = /[[:alpha:]]{2,4}/ buffer.gsub!(/(?\0') buffer.gsub!(/(?<=\s)[[:alnum:]\-_.]+@[^,<\s]+/i, '\0') Totally untested, but at least it’s somewhat easier to understand and a bit more correct. There are better ways to extract URLs and email addresses from an input than this, mind you, nikolai -- Nikolai Weibull: now available free of charge at http://bitwi.se/! Born in Chicago, IL USA; currently residing in Gothenburg, Sweden. main(){printf(&linux["\021%six\012\0"],(linux)["have"]+"fun"-97);}