From: "Peña, Botp" Date: 2008-06-23T10:32:51+09:00 Subject: Re: Regex - Exclude Multiple Characters and Global Scanning From: Ben Woodcroft [mailto:donttrustben@gmail.com] # David A. Black wrote: # > irb(main):029:0> offsets = [] # > => [] # > irb(main):030:0> str.scan(/ll./) { offsets << $~.offset(0)[1] } # > => "llgllallo" # > irb(main):031:0> offsets # > => [3, 6, 9] # That will work, thanks. It would seem intuitive to me that scan (or a # method like it) would iterate of MatchData objects, but $~ is MatchData you could wrap dBlack's hint if you want something similar to #scan eg, class String def mapscan pattern atemp=[] scan(pattern){ atemp << yield($~)} atemp end end #=> nil s #=> "llgllallo" s.mapscan(/ll./){|md| [md[0],md.offset(0)]} #=> [["llg", [0, 3]], ["lla", [3, 6]], ["llo", [6, 9]]] # # >> --------------Question 2----------------------- # >> Now an old gap in my regex understanding. How do I exclude on # >> consecutive characters? I want something like [^abc], # except aba or bbc # >> is ok, just not 'abc'. Summing this up: # > # > [^abc] means: match one character that is not 'a', not 'b', and not # > 'c'. I don't think that's what you mean. # >> reg = /something/ # >> 'abc'.match(reg) #-> no match # >> 'cba'.match(reg) #-> match # >> And then I want to be able to do OR operations too, like # not 'abc' and # >> not 'bbc', but that is probably another step of complexity. # > You can use (?!), which is negative lookahead. # > irb(main):033:0> reg = /(?!abc)[abc]{3}/ # > => /(?!abc)[abc]{3}/ # > So that means: three of a, b, c, as long as we're not looking at # > "abc" when we start looking for those three characters. # > irb(main):034:0> reg.match("abc") # > => nil # > irb(main):035:0> reg.match("abb") # > => # # > irb(main):036:0> reg.match("cba") # > => # # That is exactly what I meant. I was unaware of the negative lookahead # operator. Thanks! if you want to compare sequences, you can create a complete sequence for your case, so you do not end up creating many regex pattern. and then test everything from there. eg, SEQALPHA=("a".."z").to_a.join #=> "abcdefghijklmnopqrstuvwxyz" SEQALPHA.match "abc" #=> # SEQALPHA.match "def" #=> # SEQALPHA.match "xyz" #=> # SEQALPHA.match "bac" #=> nil SEQALPHA.match "cba" #=> nil SEQALPHA.match "yyy" #=> nil negating it on your case is simple, not SEQALPHA.match "bac" #=> true not SEQALPHA.match "abc" #=> false now using mapscan above, you can do, SEQALPHA.mapscan(/abc|xyz/){|md| [md[0],md.offset(0)]} #=> [["abc", [0, 3]], ["xyz", [23, 26]]] btw, index is a faster if you just want simple string compar. SEQALPHA.index "abc" #=> 0 SEQALPHA.index "def" #=> 3 SEQALPHA.index "xxx" #=> nil SEQALPHA.index "bac" #=> nil SEQALPHA.index /abc/ #=> 0 SEQALPHA.index /def/ #=> 3 SEQALPHA.index /efd/ #=> nil again, negating is simple not SEQALPHA.index /def/ #=> false not SEQALPHA.index /fde/ #=> true hth. kind regards -botp