From: Ammar Ali Date: 2010-11-23T06:06:44+09:00 Subject: Re: gsub and backslashes On Mon, Nov 22, 2010 at 9:25 PM, Robert Klemme wrote: > On 22.11.2010 18:21, Ammar Ali wrote: >> I don't think this is tolerance from the string parser, it is >> recognition of the \1 as a valid octal value. >> >>> irb(main):014:0>  '\1' >>> =>  "\\1" >>> irb(main):015:0>  '\\1' >>> =>  "\\1" ----8<---- > > Apparently I was not clear enough.  The point is, that there is some > tolerance.  Both sequences (line 14 and 15) produce the *same* output > although they differ in backslash usage.  This does not work if you try to > write '\' to get a single backslash.  For that you need '\\'.  If you use > two backslashes in both cases it's clear what happens and there is no room > for errors. I guess I took issue with the word tolerance. I don't think of lexers and parsers as tolerant. They are quite ruthless and dictatorial. It's either their way, or their way in a way one did not expect. :) > This is not an "escaped newline" but merely a backslash followed by > character "n".  Whether that is considered "escaped" in some way depends on > the code that processes this string.  If at all this is an escaped "n". :) You are correct sir. For someone who was nitpicking, I misspoke. :) > My precise reason to make it a rule is that it is simple and beginners do > not have to remember all these special cases that you find so worthy > mentioning. This might be six of one, half a dozen of the other kind of situation. People would start to ask if the backslash in the \n case would count in the "just add 4" rule, or not? 4 backslashes in total or 5? It seems to only shift the issue slightly, and temporarily, until one has to actually understand what is really going on. > Actually I do not like those special cases and would rather suggest to > remove them since they make things unnecessary complicated.  The repeated > occurrence of newbie confusion and the very discussion we are having here > proves that the logic creates more confusion than clarity. The only reason I > do not suggest to change this is the fact that this might break a lot of > code. I agree, but this long "heritage" that goes back to the 60s is probably very hard to shake. Maybe a new language can break away from it. Out of curiosity, what could these beasts be replaced with? Constants? Cheers, Ammar