From: Robert Klemme Date: 2007-11-03T05:45:15+09:00 Subject: Re: Database speed issues On 02.11.2007 20:43, dima wrote: > On Nov 2, 8:14 pm, rabbitb...@gmail.com wrote: >> Jeremy Woertink wrote: >>> So, I have a bit of a design problem. I have an application that work, >>> but not as well as I would like it to. My problem is that I had to >>> write an application that checks millions of records against hundreds >>> of millions of records to see if there are any duplicates. The only >>> way I could think to do this is pretty much select from the different >>> tables where the specific column matches my search string. This is >>> really slow! Obviously if it returns something, then there is a dup, >>> if not, then YAY! >>> I know there has to be a better way then just doing a SQL select >>> statement. Any ideas? >> I dunno how fast it is compared to your select statement, but if you can >> build arrays out of the data you want to compare you call the % method >> on them. Example... >> >> irb> [ 5, 10 ] & [ 10, 15 ] >> => [10] >> >> It works with more complex arrays, too... >> >> irb> [ [ 5, 10 ], [ 15, 20 ] ] & [ [ 10, 15 ], [ 5, 10] ] >> => [[5, 10]] >> >> Hope that helps... >> -- >> Posted viahttp://www.ruby-forum.com/. > > Could you be more specific? The information given is really weak. > First if your data is in some kind of SQL database first you should > look in how well you explore the power of SQL. > > In my experience if you generate your indexes right, specially the > multiple columns indexes that are in favor of your select statement > you can gain some performance boost. > > The second thing can be helping SQL database to find the right way to > do the JOINs. > This has proven mostly right for MySQL. > > If you give us some more details perhaps we can give you some more > meaningful response. Completely agree to all stated above! This should be offloaded to the database. Kind regards robert