From: Phillip Gawlowski Date: 2009-12-24T22:59:55+09:00 Subject: Re: Suggestions for a distributed job queue On 24.12.2009 03:36, Tony Arcieri wrote: > On Wed, Dec 23, 2009 at 7:17 PM, Phillip Gawlowski wrote: > >> Same thing as when you have an EC2 outage: You save state, and resume >> processing ASAP. >> > > Oof. Well for one, due to the nature of the synchronous filesystem API, > it's hard for processes in userspace to detect when things are amiss in the > underlying NFS layers. And implementing some sort of keep-alive/heartbeat system is too much work, as well. > Also, if you read my OP, saving state (aside from the state of "what jobs > have not been run yet") and recovering jobs in flight is something I want to > avoid. Well, I didn't *mean* that you implement a whole synchronization framework (what it comes down to). Alas, I implied it. > If the system fails I'd rather it simply fail and restore it to a > clean state. That way, you can have only one stateful part of the system, > and that's the only part you need to worry about recovering state from after > a failure. And this would probably be best done on "your" end of the network, too. That way you could ignore the EC2 nodes for the time being, in case of some form of network outage (assuming I understood you correctly, in that only part of your nodes are in Amazon's cloud). Not having had a look at RQueue's implemtation details, maybe you could, without too much effort, port to the Devil From Redmond's SMB system, via Samba. Or roll your own, as has been suggested, with DRb. :) -- Phillip Gawlowski Wishing everyone a merry Christmas & happy holidays. :D