From: Peter Szinek Date: 2007-06-15T15:13:21+09:00 Subject: Re: str.scan Colin, But my regexp > is rusty and I'm not sure how I would set it up > items = page.scan('
  • *
  • ') > something like that? Then items would be an array of the text in the items? Yes, they will be. However, first things first: 1) items = page.scan('
  • *
  • ') I believe you want instead is items = page.scan('
  • .*
  • ') ( or maybe items = page.scan('
  • .+
  • ') if you are not interested in empty
  • s) 2) What I really believe you want is items = page.scan('
  • .*?
  • ') ? adds greediness to your regexp - so instead of matching the first
  • . then matching as much as possible of anything, then matching the *last*
  • , 2) will match as less as possible. Let's try: stuff = <aaa
  • bbb
  • HTML >> stuff.scan(/
  • .*?<\/li>/) => ["
  • aaa
  • ", "
  • bbb
  • "] 3) Maybe you want even this: >> stuff.scan(/
  • (.*?)<\/li>/) => [["aaa"], ["bbb"]] or, even more friendly: >> stuff.scan(/
  • (.*?)<\/li>/).flatten => ["aaa", "bbb"] HTH, Peter _ http://www.rubyrailways.com :: Ruby and Web2.0 blog http://scrubyt.org :: Ruby web scraping framework http://rubykitchensink.ca/ :: The indexed archive of all things Ruby.