From: daz Date: 2005-06-26T18:40:39+09:00 Subject: Re: Best way to parse/update HTML file? Bucco wrote: > Sorry for the newbie question. This has been answered once or twice before by this group ;-) > I am trying to find the best metod for parsing a HTML file > and changinf one tag/item. Unfortunately, REXML chokes on > the file because of the incomplete tags. Completing the tag > is not an option either. What is the best way to find a specific > tag in an html file, change it's text and attribute settings? > > Thanks:) > > SA > The best way tends to involve using a package although you /could/ work your way through it using regular expressions. If you're likely to be doing this kind of thing in the future, you'll be glad you spent a bit of time installing; then it's always available. As I recall, a different package is often recommended but I don't know which is best. This is what some of us use: http://ruby-htmltools.rubyforge.org/ (Ned Konz +) Examples are included but here's another ... #----------------------------------------------------------------- EXAMPLE = < Page title
Enter a URL to make into a xxxx:

Contact - Downloads - Create - Terms of Use - List - Preferences EOX require 'html/tree' # http://ruby-htmltools.rubyforge.org/ verbose = true exa = HTMLTree::Parser.new(verbose, !false) #exa.parse_file_named('xxxx_net.html') exa.feed(EXAMPLE) # replaces '.parse_file_named' item_a = exa.html.select {|ea| ea.tag == 'a'} item_a.each {|ea| p [:ahref, ea['href']]} puts '+'*100 exa.html.each do |ea| p [ea.tag, ea['href']] ea.each do |item| if item.data? p [:data, item.to_s] elsif item.tag == 'a' item['href'].sub!(/xxxx/, 'mysite') end end puts '='*100 ea.dump end ### exa.html.dump #----------------------------------------------------------------- Output from the script above is too long to post here, so I've uploaded it to: http://www.d10.karoo.net/ruby/example_html_parse.txt Hope this is of some use, daz -- JARH (Nihon-style) http://qurl.net/h3