From: Robert Klemme Date: 2012-06-13T18:35:07+09:00 Subject: Re: Remove HTML from String? --> a better method On Tue, Jun 12, 2012 at 4:39 AM, Daniel P. C. wrote: > I hate regex.  I've written some ruby functions to remove html tags in > blocks and not just special characters... also rules for swapping html > code for anything else is included.  Example
will be swapped out > for \n with existing rules.  My code is available at > https://github.com/6ftDan/regex-is-evil What a mess. This is extremely inefficient. You create new strings all the time. You go over the string multiple times. You do not pass start and end index down to strip_seq(). There is no test which ensures start index is lower than end index (try with string ">foo<"). I'd prefer a regexp solution anytime. It's likely faster and easier to read - for me at least. Btw. /x goes a long way at making a regexp more readable - you can even include comments. Just a simple example: https://gist.github.com/2923072 But proper tool is of course a HTML parser like Nokogiri. Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/