From: Robert Klemme Date: 2010-11-23T18:17:48+09:00 Subject: Re: gsub and backslashes On Mon, Nov 22, 2010 at 10:06 PM, Ammar Ali wrote: > On Mon, Nov 22, 2010 at 9:25 PM, Robert Klemme > wrote: >> On 22.11.2010 18:21, Ammar Ali wrote: >>> I don't think this is tolerance from the string parser, it is >>> recognition of the \1 as a valid octal value. >>> >>>> irb(main):014:0>  '\1' >>>> =>  "\\1" >>>> irb(main):015:0>  '\\1' >>>> =>  "\\1" > ----8<---- >> >> Apparently I was not clear enough.  The point is, that there is some >> tolerance.  Both sequences (line 14 and 15) produce the *same* output >> although they differ in backslash usage.  This does not work if you try to >> write '\' to get a single backslash.  For that you need '\\'.  If you use >> two backslashes in both cases it's clear what happens and there is no room >> for errors. > > I guess I took issue with the word tolerance. I don't think of lexers > and parsers as tolerant. They are quite ruthless and dictatorial. It's > either their way, or their way in a way one did not expect. :) :-) But rules can be made to allow for some flexibility (just think of method calls with or without brackets in Ruby). >> This is not an "escaped newline" but merely a backslash followed by >> character "n".  Whether that is considered "escaped" in some way depends on >> the code that processes this string.  If at all this is an escaped "n". :) > > You are correct sir. For someone who was nitpicking, I misspoke. :) No problem. Apparently we both enjoy nitpicking. :-)) >> My precise reason to make it a rule is that it is simple and beginners do >> not have to remember all these special cases that you find so worthy >> mentioning. > > This might be six of one, half a dozen of the other kind of situation. > People would start to ask if the backslash in the \n case would count > in the "just add 4" rule, or not? 4 backslashes in total or 5? It > seems to only shift the issue slightly, and temporarily, until one has > to actually understand what is really going on. Hmm... Maybe. >> Actually I do not like those special cases and would rather suggest to >> remove them since they make things unnecessary complicated.  The repeated >> occurrence of newbie confusion and the very discussion we are having here >> proves that the logic creates more confusion than clarity. The only reason I >> do not suggest to change this is the fact that this might break a lot of >> code. > > I agree, but this long "heritage" that goes back to the 60s is > probably very hard to shake. Maybe a new language can break away from > it. In Ruby's case the heritage does not go back to the sixties but rather to the nineties (1997) if I am not mistaken. > Out of curiosity, what could these beasts be replaced with? Constants? I'd leave everything as is except drop special cases like '\1' (this would either be an octal escape as in a double quoted string or rather just "1"). In single quoted strings only ' would be special if preceded by a backslash. In double quoted strings I would have those characters which are special currently (", n, r, a, t and probably others I'm not thinking of right now). I am undecided whether I would make all others errors or tolerant (e.g. "\z" would either by a syntax error or just "z"). I have a slight tendency to the more strict variant though because otherwise people might be left wondering what \z means when it is just "z"; also, this would help detect typing errors (maybe someone wanted to type "\t" which is just a key away in my German keyboard). Kind regards robert -- remember.guy do |as, often| as.you_can - without end http://blog.rubybestpractices.com/