From: Peter Szinek Date: 2006-10-11T21:23:15+09:00 Subject: Re: Script to fetch Wikipedia text Hi, Well, in my opinion it is still far easier to work on the exported page text (i.e. remove MediaWiki 'markup' from there) than to scrape HTML. Tools like scrAPI (or any reasonably working web extraction tool I know of) are great if your input page is machine-generated and more or less regular. However, a wiki page does not fit into this category at all, since it is not generated and (can be) totally irregular. For some pages you could do well with HTML scraping (maybe even a lot of them), but IMHO it would be *very* hard to develop a really generic solution, working for, say 99% of the pages. (Unless all needed text blocks of wikipedia can be queried wit the same CSS selector, and no other blocks of HTML would be matched - this is possible, but quite unlikely). OTOH, the exported text is not that hairy - for the first sight you would not need more than 10 regexes to clean up the 'markup' (like [[ ]], {{ }} etc). I did not say it is a trivial task, but once you do it it will always will work properly - whereas with the HTML scraping you can be never sure. Or did I misunderstand the problem? Just to make sure: I am not speaking against scrAPI (which is cool) or web extraction (I am working and researching Web extraction on a daily basis) - but from what I understood I think it is easier to mine the exported text in this case. Just my 2c. Peter http://www.rubyrailways.com Zouplaz wrote: > Hi, > > forget regexp and use a dedicated tool like > http://rubyforge.org/projects/scrapi/ > > I've successfully build several scraper robots with it and it works just > fine ! The only drawback is too little documentation, at least for me ! > >