From: Dave Thomas Date: 2001-01-03T03:09:51+09:00 Subject: [ruby-talk:8528] Re: speedup of anagram finder jmichel@schur.institut.math.jussieu.fr (Jean Michel) writes: > Here again a good solution would be to take in account differences in > the various letters behaviour and a 'mixed basis' representation: > assume that a,e,i can occur 4 times in a word but any other letter > only 3. Then you may encode words by a polynomial of the form > a+e*5+i*5**2+b*4*5**3+etc.. (where a,e,i,b etc.. is the number of > occurences of a,e,i,b etc.. in the word). You would compute this by > pre-computing the 'powers' like 4*5**3 in the mixed basis in some array > 'pow', say, then do something like: That's a good approach. I was wondering if there's a totally different way of doing this. Right now, all solutions (including the original based on Perl) generated a hash from each word, such that hash(A) == hash(B) => A == B. How about going statistical? Anyone fancy writing a Bloom filter based finder (perhaps using #sum as the hash)? Let's think more radically! Dave