From: "M. Edward (Ed) Borasky" Date: 2007-10-28T15:48:11+09:00 Subject: Re: A memcached-like server in Ruby - feasible? ara.t.howard wrote: > > On Oct 27, 2007, at 4:31 PM, Tom Machinski wrote: > >> Hi group, >> >> I'm running a very high-load website done in Rails. >> >> The number and duration of queries per-page is killing us. So we're >> thinking of using a caching layer like memcached. Except we'd like >> something more sophisticated than memcached. >> >> Allow me to explain. >> >> memcached is like an object, with a very limited API: basically >> #get_value_by_key and #set_value_by_key. >> >> One thing we need, that isn't supported by memcached, is to be able to >> store a large set of very large objects, and then retrieve only a few >> of them by certain parameters. For example, we may want to store 100K >> Foo instances, and retrieve only the first 20 - sorted by their >> #created_on attribute - whose #bar attribute equal 23. >> >> We could store all those 100K Foo instances normally on the memcached >> server, and let the Rails process retrieve them on each request. Then >> the process could perform the filtering itself. Problem is that it's >> very suboptimal, because we'd have to transfer a lot of data to each >> process on each request, and very little of that data is actually >> needed after the processing. I.e. we would pass 100K large objects, >> while the process only really needs 20 of them. > > > i'm reading this as > > - need query > - need readonly > - need sorting > - need fast > - need server > > and thinking: how isn't this a readonly slave database? i think that > mysql can either do this with a readonly slave *or* it cannot be done > with modest resources. > > my 2cts. Add "large set of very large (binary?) objects". So ... yes, at least *one* database/server. This is exactly the sort of thing you *can* throw hardware at. I guess I'd pick PostgreSQL over MySQL for something like that, but unless you're a billionaire, I'd be doing it from disk and not from RAM. RAM-based "databases" look really attractive on paper, but they tend to look better than they really are for a lot of reasons: 1. *Good* RAM -- the kind that doesn't fall over in a ragged heap when challenged with "memtest86" -- is not inexpensive. Let's say the objects are "very large" -- how about a typical CD length of 700 MB? OK ... too big -- how about a three minute video highly compressed. How big are those puppies? Let's assume a megabyte. 100K of those is 100 GB. Wanna price 100 GB of *good* RAM? Even with compression, it doesn't take much stuff to fill up a 160 GB iPod, right? 2. A good RDBMS design / query planner is amazingly intelligent, and you can give it hints. It might take you a couple of weeks to build your indexes but your queries will run fast afterwards. 3. RAID 10 is your friend. Mirroring preserves your data when a disk dies, and striping makes it come into RAM quickly. 4. Enterprise-grade SANs have lots of buffering built in. And for that stuff, you don't have to be a billionaire -- just a plain old millionaire. "Premature optimization is the root of all evil?" Bullshit! :)