From: Bill Kelly Date: 2004-04-14T07:56:12+09:00 Subject: Re: New Ruby questions... From: "Jeff Massung" > > I've just started Ruby a couple days ago (man this is cool). Coming from > the embedded world of Forth and C, being able to do some string parsing > easily is what I'm looking forward to. So, onto a couple questions that I > can't seem to find the answers to: 26952 DUP 255 AND EMIT 256 / EMIT ( :-) > Quickly, the "teach myself" project is a simple ARM assembler. I've done > it numerous times, and it is something I'm familiar and comfortable with > (let alone something I need ATM). > > 1. How can a regexp get the longer of two possibilities that are > ambiguous? For example, I need to be able to strip out register names: > > registers = '(r0|r1|r2|r3...|r10|r11|r12...)' > > In the regular expression, how can I get it to find r10 or r11 instead of > stopping at r1? I get the same problem with opcode mnemonics (like B > instead of BX or BL). Try putting the longer strings first in the alternation sequence... irb --simple-prompt >> puts $& if "r11" =~ /r1|r2|r11/ # shorter strings first r1 >> puts $& if "r11" =~ /r11|r2|r1/ # longer strings first r11 Alternately, you could use lookahead assertions to specify the character following the token is some kind of token separating character, like... This "\b" insists a "word boundary" occurs at that point in the pattern: >> puts $& if "r11" =~ /(r1|r2|r11)\b/ r11 I think the "\b" is essentially equivalent to: >> puts $& if "r11" =~ /(r1|r2|r11)(?![A-Za-z0-9_])/ r11 Regards, Bill