From: Peter Szinek Date: 2006-04-17T08:06:42+09:00 Subject: Re: How extract data from a web site? > That depends on what data you are after, and where you want to look for it. > > If, for example, you just want to get a list of css files referenced in > a page, then regexen would likely be simpler and faster than the > tidy-up approach. Sure. But the original poster mentioned some HTML parsing so i have thought regexps are not enough. > Sort of. I've seen tidy make some odd assumptions about what the > "correct" output should be, based on surreal HTML input. And this can > throw off the XML manipulation code. Of course. tidy is just 'better than nothing'. I did not mean it will work everywhere (it certainly won't) - but at least you can get closer to your goal (in some cases) > There is a real cost to making the parsing/cleaning code highly robust, > and if you can make certain assumptions about the source text (and live > with the risks that things can change), you can often make the app > faster/simpler. Well, for a *really* robust something, take a look at my earlier mail in this thread. You can not get nowhere near to that with any other tool/technique (if you think yes, LMK). I am in the web extraction business, usually we are extracting data from hundreds of thousands of pages on a daily basis so i have some experience with this stuff. Our wraper generator solutions are usable for, well, most of the pages out there (say 95%), utilizing adaptive techniques if the page changes and other stuff for robustness etc. But this software took 5 years to develop for a medium-sized team, and now that it is finished we can almost rewrite it from sracth because it is nearly unusable on some of the web2.0 pages... (due to AJAX etc) cheers, Peter