From: Vincent Isambart Date: 2007-07-02T21:38:23+09:00 Subject: [PATCH] Problem with ruby 1.8.6-p36 (and p39) on Tiger ------=_Part_58594_14597293.1183379905281 Content-Type: text/plain; charset=ISO-8859-1; format=flowed Content-Transfer-Encoding: 7bit Content-Disposition: inline Hi, > The problem can be easily reproduced under Tiger with a one-liner: > % ruby -ve 'require "timeout" ; Timeout.timeout(5) { puts `df -h` }' > ruby 1.8.6 (2007-06-07 patchlevel 36) [i686-darwin8.10.1] > /usr/local/lib/ruby/1.8/timeout.rb:54: execution expired (Timeout::Error) > from /usr/local/lib/ruby/1.8/timeout.rb:56:in `timeout' > from -e:1 I investigated more and the function that blocks is rb_thread_cancel_timer (eval.c). This function is called in proc_exec_v (process.h) only in 1.8.6-p36 and -p39, not in the 1.8 branch nor in 1.8.6-p0. And it is empty if pthread is not enabled. So I investigated a little more to see what was wrong in this rb_thread_cancel_timer function, and why everything works fine under systems other than Tiger (Linux, Leopard...). The reason seems to be due to a bug in Tiger. Here is a little step-by-step of what happens when calling in Ruby system() or ``: - the process is forked, the parent does not do much but the child do what follows ; - the rb_thread_cancel_timer is called to kill the timer thread. However, as this process is a child created using fork, it has no timer thread (fork kills all threads in the child process except the current one), and the time_thread references a thread in the parent. So on most systems, the pthread_cancel and pthread_join done by rb_thread_cancel_timer both fail. But under Tiger pthread_cancel SUCCEEDS and pthread_join HANGS. However, after seeing in the ChangeLog why this call to rb_thread_cancel_timer is done (cf ruby-dev:30581), it's clear that this cannot be simply removed. However it should not be called just in a child created by fork in C. The simplest way may be to just put time_thread_alive_p to false in the children of forks created to start a process (for example in pipe_open in io.c or in rb_f_system in process.c). However I do not know well enough Ruby to be 100 % sure of this. I attached a small patch using pthread_atfork to do this and it solves at least on my small tests the problem with system() or `` in Timeout.timeout, and the problem mentioned in ruby-dev:30581 is still corrected. But I think it needs some reviews to be sure it does not add any problem. Cheers, Vincent Isambart ------=_Part_58594_14597293.1183379905281 Content-Type: application/octet-stream; name=thread_not_alive_anymore.patch Content-Transfer-Encoding: base64 X-Attachment-Id: f_f3mxlebh Content-Disposition: attachment; filename="thread_not_alive_anymore.patch" SW5kZXg6IGV2YWwuYwo9PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09PT09 PT09PT09PT09PT09PT09PT09PT09PT09Ci0tLSBldmFsLmMJKHJldmlzaW9uIDEyNjc5KQorKysg ZXZhbC5jCSh3b3JraW5nIGNvcHkpCkBAIC0xMTgyNiw2ICsxMTgyNiwxMiBAQAogewogfQogCitz dGF0aWMgdm9pZAorcmJfY2hpbGRfYXRfZm9yayh2b2lkKQoreworICAgIHRpbWVfdGhyZWFkX2Fs aXZlX3AgPSAwOworfQorCiB2b2lkCiByYl90aHJlYWRfY2FuY2VsX3RpbWVyKCkKIHsKQEAgLTEx OTIwLDYgKzExOTI2LDcgQEAKICNpZmRlZiBfVEhSRUFEX1NBRkUKIAlwdGhyZWFkX2NyZWF0ZSgm dGltZV90aHJlYWQsIDAsIHRocmVhZF90aW1lciwgMCk7CiAgICAgICAgIHRpbWVfdGhyZWFkX2Fs aXZlX3AgPSAxOworICAgICAgICBwdGhyZWFkX2F0Zm9yayhOVUxMLCBOVUxMLCByYl9jaGlsZF9h dF9mb3JrKTsKICNlbHNlCiAJcmJfdGhyZWFkX3N0YXJ0X3RpbWVyKCk7CiAjZW5kaWYK ------=_Part_58594_14597293.1183379905281--