From: Simon Krahnke Date: 2009-05-06T16:05:03+09:00 Subject: Re: removing Whitespace using regexp * Arun Kumar (07:54) schrieb: > Hi, > Previously I posted a topic on how to strip all html tags and getting > the remaining text using regexp. Luckily I got one. This is the regexp: > > /([^>]*)(?=<[^>]*?>)/im And what do you do with this regexp? > In this case I'm able to get all the data between the html tags. But one > small problem. Hasn't everybody told you, there are problems with parsing HTML with regexps? > This is the output which I get when I parse the html content of > example.com using the above regexp. Here you can see some white space > between the data(ie. between 'Example web page' and 'You have > reached...'. These whitespaces are generated in place of the html tags > which I avoided using the above regexp. Really? Aren't they just from all the meaningless whitespace that's in a typical HTML document? > I want to remove those > whitespaces. I think that modifying the above regexp will give me the > right output without white spaces. Can somebody please help me. There are easy ways to strip all the whitespace, which is certainly not what you want, and there is a simple way to reduce all runs of whitespace by just one space (gsub(/\s+/, ' '), which probably also not what you want. Selectively removing some of the whitespace isn't easy at all, but it is probably a lot easier with a real HTML parser. mfg, simon .... l