From: James Gray Date: 2009-07-07T07:43:14+09:00 Subject: Re: file.seek and unused bytes On Jul 6, 2009, at 5:08 PM, Greg Willits wrote: > In my original system, I used the row lengths and the number of rows > per > "set" to pre-calculate a lookup of offsets into the start of each set. > Upon recieving a request for record 623,456 I first had to determine > which set that would be in (4th set based on 200,000 per set). > By saving the row-specific offsets as an array, that is simplified > to : > > record_start = @record_offsets[row_number] > > This turned out to make the total process of fetching rows 15% faster. > (roughly 32 seconds vs 38 seconds for reading 1.4 million rows on my > dev > system). > I already index the data itself using hashes (which work very, very > fast) for aggregation lookups, so this concept is quite parallel to > other ways the code works, and is a worthwhile change to make -- so > thanks for that idea. The more I read in this thread, the more I feel it's crying out for a database. I recognize that you disagree with this, I just disagree with your disagreement. In a totally friendly way, of course. :) Originally I too was thinking of the already suggested SQLite, but now I'm thinking that not right. My current feeling is that Redis would be a terrific fit: http://code.google.com/p/redis/ That's just my two bits. I promise to let the database argument lie now. James Edward Gray II