From: Ken Bloom Date: 2006-11-13T23:50:05+09:00 Subject: Re: regular expression too big On Mon, 13 Nov 2006 00:07:52 -0800, Paul Lutus wrote: > Peter Schrammel wrote: > >>> When you post an question like this, it is always a good idea to >> reveal the >>> purpose as well as the method. >>> >>> >> >> First of all: Thanks for the replies, I think I have enough input to >> chew on. >> >> And to reveal the purpose: >> I'd like to match a LOT of words/strings (words with spaces) and >> sometimes regexes against a lot of small and medium size( < 10000 byte ) >> strings. > > If you are going to compare strings to strings as well as strings to > regexes, maybe a two-tiered scheme would be better. First tier, simple > textual comparison using a hash of strings. Second tier, regexes. > > This removes the ambiguity that a particular data string might pass a string > comparison but fail any of the provided regexes, or the opposite. It will > also be much faster than a gigantic list of strings and regexes, all > presented to the regex engine as though they were all regular expressions, > even though some are simple string comparisons. > > I think this (e.g speed improvement) would have been true in Perl also. > His regex isn't anchored to the beginning and end of the string. This makes hashes useless for the kind of comparison he seems to want to perform. --Ken -- Ken Bloom. PhD candidate. Linguistic Cognition Laboratory. Department of Computer Science. Illinois Institute of Technology. http://www.iit.edu/~kbloom1/