From: Robert Dober Date: 2006-05-24T04:53:14+09:00 Subject: Re: Regexp help: Matching HTML having trouble w/greediness ------=_Part_163313_2785522.1148413989305 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: quoted-printable Content-Disposition: inline On 5/23/06, Wes Gamble wrote: > > Wes Gamble wrote: > > All, > > > > I am attempting to do some matching on some HTML. > > > > Here's what I want: > > > > I want to be able to match any tag which is of the form > ...> that does NOT contain a "mailto:" href. The grouping is for some > > substitution that I'm doing. > > > > Here's my pattern: > > > > /()/mi > > > > What I find is that this pattern will successfully handle most area > > tags, however, when an area tag is followed by a tag (which also ha= s > > a href attribute, it will match everything between > the tag). So, for example, > > > > other stuff, including tags > href=3D"blah"> > > > > this pattern matches all the way through the end of the tag. > > > > I tried to stick a negative lookahead (?!<) at the end of the pattern > > but that doesn't seem to help. > > > > How do I get this pattern to STOP matching? > > > > Thanks for any help, > > Wes > > > This appears to work better: > > /(]*?)(href=3D['|"])(?!mailto:)(.*?)(['|"].*?>)/mi > > The "any character but '>'" seems to stop the evaluation of the match > from making it any further than the end of the tag. > > Wes > > -- > Posted via http://www.ruby-forum.com/. > > So I was just typing and typing and typing... Glad Ur happy with it. Robert --=20 Deux choses sont infinies : l'univers et la b=EAtise humaine ; en ce qui concerne l'univers, je n'en ai pas acquis la certitude absolue. - Albert Einstein ------=_Part_163313_2785522.1148413989305--