From: nobuyoshi nakada Date: 2005-05-20T16:01:01+09:00 Subject: Re: [ ruby-Bugs-1930 ] CGI::escapeHTML escapes the ampersand in an existing escape command Hi, At Thu, 19 May 2005 20:55:50 +0900, Zev Blut wrote in [ruby-core:05046]: > >> > The way to let escapeHTML not to escape a particular ampersand is > >> > another story. > >> > >> I can see the need for the current behavior and of course I have the > >> need for the changed behavior. Two techniques for providing a > >> preserving escape come to my mind. One would be to add an optional > >> boolean argument that when true will keep escape commands, the > >> default value is false and thus current usage will not change. Another > >> is to add another method that implements the patch provided in the > >> report. > > > > Leaving ampersands breaks round-trip by escapeHTML and unescapeHTML. > > I think it would be a problem. > > I don't think this breaks round-trip usage. See below: Rather I'm afread about < and > especialy. They may lead security problem. But I noticed that CGI.unescapeHTML doesn't replace other than amp, quot, lt and gt, but it leaves euro, yen and so on. Finally, it may not be a problem except for those HTML special characters. Index: lib/cgi.rb =================================================================== RCS file: /cvs/ruby/src/ruby/lib/cgi.rb,v retrieving revision 1.79 diff -U2 -p -r1.79 cgi.rb --- lib/cgi.rb 1 Nov 2004 23:49:15 -0000 1.79 +++ lib/cgi.rb 20 May 2005 05:47:01 -0000 @@ -320,4 +320,7 @@ class CGI RFC822_MONTHS = %w[ Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec ] + # Ampersands escaping HTML special characters. + AMPERSAND = /&(?=(?:amp|quot|[gl]t|\#(?:x0*(?:2[26]|3[ce])|0*(?:3[48]|6[02])));|(?!\w+;))/ni + # :startdoc: @@ -360,5 +363,5 @@ class CGI # # => "Usage: foo "bar" <baz>" def CGI::escapeHTML(string) - string.gsub(/&/n, '&').gsub(/\"/n, '"').gsub(/>/n, '>').gsub(//n, '>').gsub(/