From: Peter Schrammel Date: 2006-11-18T03:15:07+09:00 Subject: Re: regular expression too big Paul Lutus wrote: > Peter Schrammel wrote: > >> Paul Lutus wrote: >>> Yes, unless he is matching whole words, he is stuck with regexes. There >>> is very likely to be a refactoring for this problem, and it would have to >>> start with a clear statement of the problem to be solved. >>> >> Sorry, my fault. >> >> The problem is to match a whole bunch (>70000) of words (later regexps) >> on a string. > > You are not being very clear. Do you mean to match entire words, letter for > letter, beginning to end, at least some of the time? If so, for those > cases, you can use a hash table that is preloaded with the words as keys. > That will be fast. no its a substring match kind of /hello/.match("dear customer id like to say hello to ...."). I have to find 70000 words in a large textfile: if lotofwords.match(bigtext) ..... I need a trigger to classify the bitext depending on wether any of the word (later regexps) are in it or not. Tries work well for substring matching in this case but fail with regexps. > Also precompile the regexes before use. You probably already know this, but > I thought I would mention it anyway. yes, done that. > Another option is C++, which has a readily available regexp library. The way > I would go about this is to design the entire thing in Ruby, then, if the > speed was not acceptable, recreate it in C++. This gives you the advantage > of speedy development in Ruby, followed by speedy execution. > > If this is a full-on language analysis problem, you really should be using > Prolog or Lisp anyway. If the problem really is as complex as you are > hinting at, you may not be using the right language, or even class of > language. > yes, but I wanted to try it with "my favorit" language first before doing something wild. It's a rails application so I'd have to call lisp/perl/.... programms, interfacing them (calling on each test would be a mess because reading somethousand lines every time is a performance killer.) Anyway, thanks for the replies I'll give a perl interface a try (due to some other reasons like robust html parsing ...).