From: Peter Szinek Date: 2006-11-22T20:03:54+09:00 Subject: Re: Hpricot/Rubyful Soup comparison Luciano Ramalho wrote: > On 11/22/06, Peter Szinek wrote: >> I did not do any benchmarks, but I am scraping a lot of relatively big >> pages on a daily basis and I can tell you, RubyfulSoup is magnitudes >> slower than HPricot. > > HPricot is partially written in C, so it should be faster than a > pure-Ruby lib like RubyfulSoup. true > Also, RubyfulSoup aims to be very resilient to malformed markup, So it's HPricot. HPricot is not just a HTML parser which can parse (relatively) valid HTML - it can parse any HTML 'somehow'. We can argue whether HPricot's 'somehow' is better or worse that RubyfulSoup's, but it is a fact that HPricot is handling malformed pages very well. so it > must resort to heuristics that have a performance cost. I don't know > fow HPricot handles HTML or XML with really serious flaws like tags > that open but never close and so on, This concretely is absolutely OK. Maybe we would need a list of serious problems and see how Hpricot vs RubyfulSoup is handling them. From what I have seen, HPricot did not have any problems with any page... > has managed to deal amazingly well with such problems. If you need to > parse low quality markup, the performance penalty of RubyfulSoup may > be well worth the price. I am still not sure what are the added benefits of RubyfulSoup parsing over HPricot (although I am not claiming that there are none) - I would like to see a real serious comparison to decide this... Peter __ http://www.rubyrailways.com