From: "Jesús Gabriel y Galán" Date: 2010-08-04T15:28:10+09:00 Subject: Re: Get rid of extra, blank lines via html parsing? On Wed, Aug 4, 2010 at 6:29 AM, David Ainley wrote: > So I am trying to get some information from a snippet of html > (http://pastebin.com/iTXyxQ0j), and im using doc.inner_text to get the > important parts, but when I do so I get an odd amount of spacing > (http://pastebin.com/6HWDs5dm).  is there a way where I can get rid of > all that extra spacing so I can just print the output and it looks > clean?  possibly something like > > pino > 0.2.11-ubuntu0~lucid > troorl > (2010-07-04) > > pino > 0.2.10-ubuntu0~karmic > troorl > (2010-05-27) > > that?  or can i get each piece of text and add it to an array?  if i do > that while its got all that odd spacing, is that spacing a piece of the > variable?  or is it juts the text? You can remove 2 or more consecutive "\n" like this: irb(main):001:0> s =< "test\n\ntest2\nsdfsdf\nwerwer\n\n\n\n\nsdfsdfsd\nsdfer234\n" irb(main):019:0> s.gsub /\n\n+/, "\n" => "test\ntest2\nsdfsdf\nwerwer\nsdfsdfsd\nsdfer234\n" or irb(main):020:0> s.gsub /\n{2,}/, "\n" => "test\ntest2\nsdfsdf\nwerwer\nsdfsdfsd\nsdfer234\n" Hope this helps, Jesus.