From: Chris Pine Date: 2003-04-17T00:14:58+09:00 Subject: Re: regular expressions ----- Original Message ----- From: "Mauricio Fern�ndez" It seems to me you might like this: http://mywebpages.comcast.net/mdschneider/ruby/index.html#REGTREE ---------------------------- Hmmm... I really like what he has to say on his webpage, but that is just *way* too verbose. ----- Original Message ----- From: "Daniel Carrera" If my theory of (RE) evolution is correct, ---------------------------- It seems that a mathematician (of course) invented regular expressions: Kleene. Ken Thompson then built them into qed, then ed, then grep. From there they took off. So, no, they were not intended to be used as they now are. (I mean the current syntax... it really is natural and easy in a math class!) ----- Original Message ----- From: "Arthur Chance" <{spamtrap}@qeng-ho.org> Well, a long time ago, in a very different world, there was a pattern matching language called Snobol ---------------------------- Interesting! I'll look into it. What is a regular expression? Basically, it's a terse notation for a set of (perhaps countably (infinitely) many) strings. So this has got me to thinking about what we use regexps for (since PCREs do a lot of little things you don't get in old-school regular expressions). Given a string `str' and a set of strings `S', we want to know: - Is `str' in `S'? - Is some substring of `str' in `S'? --- If so, which one? Where does it start? --- And is there another one after the end of that one? Sometimes we also want this information about the sub-regexps in a regexp (the stuff in parentheses). Then there's gsub... I'll skip that for now. It seems that the major problem is that regexps are used on strings, but they are implemented as if they *are* strings, so you have to escape everything. It's like programming in a language where `(' is a reasonable variable name! So here's a first shot (and I'll use `<' and `>' as delimiters, so as not to suggest that I am trying to get this implemented over Ruby's own `/.../' regexps; obviously this wouldn't work at all in Ruby; I'm just wondering if this is easier to read): <^ 'abc'* [^'!@#$']> which would match 'abba' and 'ccr', but not 'bab!'. Perhaps if we changed the second `^' to `!' (more natural) and twiddle a bit: [^, 'abc'*, !'!@#$'.1] Hmmm... I like the `.1', but I don't like the array-style and the commas. Well, that's a beginning, anyway. Thoughts? Chris