From: "Jesús Gabriel y Galán" Date: 2010-10-09T17:02:10+09:00 Subject: Re: How can I count number of elements in an HTML page On Fri, Oct 8, 2010 at 11:10 PM, Paul wrote: > On Oct 5, 9:33 pm, Steel Steel wrote: >> > I can find the section I want with a regex but I don't know how to >> > iterate through the string looking for particular elements.  I was >> > thinking about taking the section I'm interested in and saving it as >> > an array and then iterating through each array element (html line) >> > that way, but I thought there might be a quicker way to do it. >> >> $html.scan(%r{}m).to_s.scan(/
  • /).size > > Thanks Steel.  This worked fine.  I just needed to make it a lazy > search with .*? > > I've got nothing against Nokogiri or the other solutions but I was > hoping for a solution like this that just uses the core libraries for > portability. You have to be careful, then, about all the possible combinations that make a valid HTML but make the above regexp fail: irb(main):025:0> string=<
    • something
    • something else
    irb(main):027:0" EOS => "
    • something
    • something else
    \n" irb(main):029:0> string.scan(%r{}m).to_s.scan(/
  • /).size => 1 irb(main):031:0> require 'nokogiri' => true irb(main):032:0> doc = Nokogiri::HTML(string) => #, #] children=[#]>, #]>]>]>]>]>]> irb(main):033:0> doc.search("/html/body/div/ul/li").size => 2 In general: parsing HTML with regexp can get messy. Best leave the work to a proper library that handles all the strange nuances. Jesus.