From: Timothy Goddard Date: 2006-02-09T17:38:23+09:00 Subject: Re: Using Classifier::LSI I'm doing a project at the moment cataloguing publications where I've been using Ferret (a ruby port of lucene). I have to say, this is a brilliant library. It has really good search capablilities, division of 'documents' into various fields that can be individually searched, and it's fast. Here are some basic benchmarks using the most simple entry mode (both stores and tokenises all input) using ruby 1.8.4 (2005-10-29) [i686-linux] on a 1.4GHz Celeron M processor, 768MB memory: tim@ghostknife:~/programming/ruby> cat bm_ferret.rb require 'benchmark' require 'rubygems' require 'ferret' $index = Ferret::Index::Index.new :path => './test_index', :create => true $words = [] File.open('/usr/share/dict/words') do |dict| dict.each_line do |line| $words << line.chomp end end def get_words ret = [] 10.times do ret << $words[(rand * $words.length).to_i] end ret end Benchmark.bmbm(7) do |x| x.report('Filing: ') do 1000.times {$index << {:key => (rand * 1000000).to_i, :data => get_words.join(' ')}} end end tim@ghostknife:~/programming/ruby> ruby bm_ferret.rb Rehearsal -------------------------------------------- Filing: 7.920000 0.630000 8.550000 ( 8.551508) ----------------------------------- total: 8.550000sec user system total real Filing: 7.680000 0.700000 8.380000 ( 8.637888)