From: Bill Kelly Date: 2006-03-08T04:36:10+09:00 Subject: Re: Is there link extractor or similar html processing libs for Ruby From: "James Edward Gray II" > > > > You meant something like this ? (quite dirty but works) > > > > puts open("some.html").read.scan(//) > > No, it doesn't, trust me. ;) Toss a simple "\n" in there and you're > sunk: > > href="whatever"> > > Parsing HTML is hard and you don't want to use regular expressions to > do it. Hi, not trying to be argumentative, just surprised. I thought parsing HTML with regexps was pretty easy. Well, lexing HTML into tokens, I mean. Since there are no recursive structures (that I know of) in the syntax for an open or closing tag, it seemed reasonably well suited to regexps to me. . . . . Heheh, or maybe the passage of time has given the memories a rosy glow. I just looked up the last HTML lexer I wrote, 5 years ago, and it's 19 lines of regexp. Admittedlly it's a very clean 19 lines, but still, lengthier than I remembered.... :) Regards, Bill