From: Robert Klemme Date: 2007-02-06T07:10:13+09:00 Subject: Re: Regular expressions (extracting urls) On 05.02.2007 21:37, David Krmpotic wrote: > Hi! > > I have to extract an url from the text and make it a link (a href...).. > The trick is that I have to be careful not to replace the url, that are > already a part of the link. > > so: > > link = "Go here: http://www.something.com!" > link.gsub!(/https?:\/\/[a-z0-9\.\-\_=&\+\/\?]+/i, ' href=\'\0\'>\0') > > link becomes: > link = "Go here: href='http://www.something.com'>http://www.something.com!" > > Now... > > When the link is this: > link = "Go here: href='http://www.something.com'>http://www.something.com!" > > The regular expression must not replace it! > > I know that for example if I want to exclude the links that start with > xhttp, I can write: > link.gsub!(/([^x]https?:\/\/[a-z0-9\.\-\_=&\+\/\?]+)/i, ' href=\'\1\'>\1') > > but how can I exclude links that start with href=" and href=' ? > The problem is that I don't know how to specify that HREF cannot > preceede the link (cannot write [^href], [^(href)] doesn't seem to work > either and it also screws \n ... ) Maybe it's enough to do s.gsub( /([^"'])\b(http[^\s"']+)\b([^"'])/, '\\1\\2\\3' ) That depends on your input text. This piece has some weaknesses, e.g. won't substitute hrefs at the beginning and end of the string (you could pad with a whitespace). Kind regards robert