From: dima Date: 2007-11-03T04:45:04+09:00 Subject: Re: Database speed issues On Nov 2, 8:14 pm, rabbitb...@gmail.com wrote: > Jeremy Woertink wrote: > > So, I have a bit of a design problem. I have an application that work, > > but not as well as I would like it to. My problem is that I had to > > write an application that checks millions of records against hundreds > > of millions of records to see if there are any duplicates. The only > > way I could think to do this is pretty much select from the different > > tables where the specific column matches my search string. This is > > really slow! Obviously if it returns something, then there is a dup, > > if not, then YAY! > > > I know there has to be a better way then just doing a SQL select > > statement. Any ideas? > > I dunno how fast it is compared to your select statement, but if you can > build arrays out of the data you want to compare you call the % method > on them. Example... > > irb> [ 5, 10 ] & [ 10, 15 ] > => [10] > > It works with more complex arrays, too... > > irb> [ [ 5, 10 ], [ 15, 20 ] ] & [ [ 10, 15 ], [ 5, 10] ] > => [[5, 10]] > > Hope that helps... > -- > Posted viahttp://www.ruby-forum.com/. Could you be more specific? First if your data is in some kind of SQL database first you should look in how well you explore the power of SQL. In my experience if you generate your indexes right, specially the multiple columns indexes that are in favor of your select statement you can gain some performance boost. The second thing can be helping SQL database to find the right way to do the JOINs. This has proven mostly right for MySQL. If you give us some more details perhaps we can give you some more meaningful response.