From: G_ F_ <8si.greg@...> Date: 2009-09-08T03:05:33+09:00 Subject: Re: rss parsing error. Young Gyu Park wrote: > At these days, I try to parse 'http://www.forbes.com/news/index.xml' > using > feedzirra. > As you access this url, you can recognize what the problem is. > > They added an unnecessary html tag which made malformed rss format. Glancing at the output of their feed I see no malformed RSS. I do see them "exercising some options" that most feeds don't, such as embedding CDATA in the link tags. Using Nokogiri to parse this feed is easy: #!/usr/bin/env ruby -wKU require 'rubygems' require 'nokogiri' require 'open-uri' url = 'http://www.forbes.com/news/index.xml' xml = Nokogiri::XML(open(url)) puts "Feed title: #{ (xml%'title').content }" puts "Feed description: #{ (xml%'description').content }" puts "Feed link: #{ (xml%'link').content }" # get the first item item = (xml/'item').first puts "Item title: #{ (item%'title').content }" puts "Item link: #{ (item%'link').content }" puts "Item pubDate: #{ (item%'pubDate').content }" puts "Item description: #{ (item%'description').content }" puts "Item author: #{ (item%'author').content }" Not all feeds are this straightforward or well constructed. That's where using a pre-built library to parse comes in handy but I haven't found one yet that handles everything out there correctly. Even Google's reader gets it wrong on some malformed feeds. Aaron Patterson (AKA tenderlove) has done a great job with Nokogiri. I've tested a lot of feeds and seen occasions where the built-in RSS reader and other libraries puked or spun off and never returned. I've run into feeds that caused Hpricot to be unable to strip broken HTML embedded inside the descriptions, but Nokogiri was able to handle it. So, if you can't get a library to do what you want, jump in with Nokogiri and give it a try. -- Posted via http://www.ruby-forum.com/.