From: Ned Konz Date: 2002-05-18T00:45:08+09:00 Subject: Re: Greedy Regexp Aidan wrote: > Are Ruby regular expressions, especially those built using the > Regexp::MULTILINE option, greedy by default? > > Is there anyway to make them non greedy? > > I need to scan over a large HTML table which consists of a large number of > table row units: > > > stuff > other stuff > > > stuff > other stuff > > > and so on ... The program needs to extract each unit and process > it individually. Will a RE like > > .* > > need to recoded with some guards against greedyness? Parsing HTML with regexes is not a good idea (that is, it's more complex than you'd think). For specific files you could use something like .*? the following question mark makes it non-greedy. But you should also look at the html-parser module. I've written a module that knows about HTML structure and am writing one that builds a tree.