From: Tony Arcieri Date: 2009-12-24T11:36:05+09:00 Subject: Re: Suggestions for a distributed job queue --000e0cdf1baae327cd047b704c60 Content-Type: text/plain; charset=ISO-8859-1 On Wed, Dec 23, 2009 at 7:17 PM, Phillip Gawlowski wrote: > Same thing as when you have an EC2 outage: You save state, and resume > processing ASAP. > Oof. Well for one, due to the nature of the synchronous filesystem API, it's hard for processes in userspace to detect when things are amiss in the underlying NFS layers. Also, if you read my OP, saving state (aside from the state of "what jobs have not been run yet") and recovering jobs in flight is something I want to avoid. If the system fails I'd rather it simply fail and restore it to a clean state. That way, you can have only one stateful part of the system, and that's the only part you need to worry about recovering state from after a failure. -- Tony Arcieri Medioh! A Kudelski Brand --000e0cdf1baae327cd047b704c60--