From: Joel VanderWerf Date: 2003-01-27T05:29:27+09:00 Subject: Re: solaris porting problem -- flock failure? ahoward wrote: > On Sun, 26 Jan 2003, Joel VanderWerf wrote: > > from http://path.berkeley.edu/~vjoel/ruby/solaris-bug.rb > > # ugly hack because waiting for a lock in a Ruby thread blocks the process > > > this is the problem i think. if you look at the file.c implementation of > rb_file_flock you'll notice that it _used to_ do something very similar at the > C level - but that this has been _removed_ > > file.c:1531 #if defined(EWOULDBLOCK) && 0 > .... > > which you've essentially re-implemented in ruby. Interesting. I checked again using two irb sessions that waiting for a lock does block all ruby threads. Also, I ran my solaris-bug.rb without the polling code (the ugly hack), instead just calling #flock. I expected deadlock. Over several runs, I've seen two kinds of results: 1. solaris-bug.rb:43:in `flock': Interrupted system call - "/tmp/test-file-lock.dat" (Errno::EINTR) 2. solaris-bug.rb:128:in `load': marshal data too short (ArgumentError) The latter looks like the same collision problem as usual. Could #1 be a result of deadlock, which I would expect to happen eventually? Anyway, the fact the #2 happens (no pun intended) at all suggests that polling isn't really the cause of the problem, and blocking the whole process doesn't prevent the collision. > i'm guessing someone smarter than us realized this was not safe. thus, it > appears that the correct behavior for flock is block the entire process and > that threads should not be using flock at all. i do not understand the exact > reason for this, but searching google for flock/fcntl/thread will bring up a > plethora of problems. i think you can be farily certain that the behavior of > flock from inside a threaded program will be ill-defined across OSs. It's particularly confusing to me because ruby threads all run in the same native thread, so I expected to be insulated from any multithreading problems with locks or with anything else at the system level, as long as I was careful to manage concurrency among my threads (keep reader counters and use LOCK_NB, use Thread.critical, etc.). The test program runs perfectly with thread_count == 1 (that's ruby threads, not native threads, and it doesn't count the "supervisor thread", only the "workers", so there are actually still 2 threads). So I have trouble believing that the underlying flock implementation (in ruby source or in the depths of solaris) is bad. It seems to me that as long as ruby's threading code is correct, and I use appropriate threading constructs correctly, this should all work. > as to a fix : keeping in mind that file locks are only (usually) advisory you > don't really buy anything using flock over a mutex since any other process can > chose to clobber the file anyhow - the lock will not prevent this. this is > troublesome since you also want to fork... the only way i can think of doing > this is to have each thread ask it's parent process to lock the file in it's > stead (in a critical section), rather than attempting to do so itsef. even > this might not be safe and you might have to resort to some sort of IPC to > ensure single writer semantics. the sysvipc module from the raa seems to be > unreachable right now. does anyone have a copy? I don't mind using advisory locks, since I don't expect any processes but my own (which use the still semi-functional locking code) to access the files of interest. Also, I don't fork within the context of a lock. That workaround sounds pretty scary, and potentially less portable, but what do I know, maybe I should learn how to use inter-process mutexes...