From: Martin DeMello Date: 2008-10-07T09:17:46+09:00 Subject: Re: Problem with comparing huge amount of strings On Mon, Oct 6, 2008 at 12:12 PM, Ragav Satish wrote: > 4. There are more complicated blocking schemes - like n-gram(chunk by N > common continuous characters) or even sampling based ones. n-gram were my first thought - indeed, I'd started writing out a bigram-based scheme, then I realised that it'd fail badly if there was a common word like "systems" or "computers" that a lot of the entries had. Maybe some sort of multipass scheme to first reduce each entry to a characteristic word, then do n-gram frequency analysis on those words (my idea was this: pass 1: make up a frequency table of bigrams, pass 2: characterise each entry by the presence/multiplicity of the six or so most common ones) martin