From: "trans. (T. Onoma)" Date: 2004-11-12T01:38:55+09:00 Subject: Re: #scan with or'd (`|`) subexpressions. On Thursday 11 November 2004 10:54 am, Warren Brown wrote: | T., | | > Does the new Ruby regexp engine do this? | > | > irb(main):001:0> '1234'.scan(/(1)(2)|(3)(4)/) | > => [["1", "2", nil, nil], [nil, nil, "3", "4"]] | > irb(main):002:0> | > | > Why would all the subexpressions be listed when there | > is an `|` (or) used? | | For collecting matches, Ruby simply looks at opening parenthesis - | nothing else. The part of the string matched by the regular expression | delimited by the first open parenthesis and its matching close | parenthesis will be the first match, the second opening parenthesis and | its matching close parenthesis will define the next match, etc. I see. Perhaps there is good reason for this. But I just don't see it. IN practice it causes me to have to strip out a whole lot of nils from the results. Honestly, I can't see how it makes any sense. The regexp will match on the first "or" that succeeds, right? So all the others are by necessity nil. But perhaps I'm overlooking some possibility. | I have not yet had an opportunity to play with Oniguruma, so I can't | say definitively if behaves the same way. However I would be very | surprised if it didn't, since virtually every other language behaves | this way. Neither have I. But I do hope Oniguruma is a better than "every other". By the way, have you read about Perl 6 new RE engine? I must say it look pretty sweet. Theses are definitely not your average everyday Regular Expressions. It now allows you to create your own rules and encapsulate those and resuse them -- much more like a grammer parser. Thanks, T.