From: Paul Lutus Date: 2006-11-19T04:35:08+09:00 Subject: Re: regular expression too big Peter Schrammel wrote: / ... >>> The problem is to match a whole bunch (>70000) of words (later regexps) >>> on a string. >> >> You are not being very clear. Do you mean to match entire words, letter >> for letter, beginning to end, at least some of the time? If so, for those >> cases, you can use a hash table that is preloaded with the words as keys. >> That will be fast. > > no its a substring match kind of /hello/.match("dear customer id like to > say hello to ...."). Translation: "Yes, I am searching for matches with individual words or groups of words, character by character, without always requiring regular expressions." In this case, a hash approach is indicated for one tier of your solution. You need to realize that this: /hello/.match("dear customer id like to say hello to ....") Is fantastically, unbelievably, slower than this: hash = [ "hello" => true,"goodbye" => true,"stay a while" => true ] if hash[word] ... All you have to do is tokenize your word list using "split" or another similar approach, then apply the hash to it. You may have a need for regular expressions, but it is clear that you can also use a hash and save a load of run time. > I have to find 70000 words in a large textfile: > if lotofwords.match(bigtext) ..... > I need a trigger to classify the bitext depending on wether any of the > word (later regexps) are in it or not. Tries work well for substring > matching in this case but fail with regexps. At this point, to offer any useful advice, it would be nice to know what the goal is, more specifically than we've heard up to now. / ... > Anyway, thanks for the replies I'll give a perl interface a try (due to > some other reasons like robust html parsing ...). Oh, Ruby can parse HTML just fine, and in only a few lines. -- Paul Lutus http://www.arachnoid.com