From: "Marshall T. Vandegrift" Date: 2006-08-28T02:51:18+09:00 Subject: Re: Can Anyone Explain This Memory Leak? Zed Shaw writes: > I boiled the problem down to this: > > http://pastie.caboo.se/10194 And this one, from the other thread: http://pastie.caboo.se/10317 > It's a graph of the "leak" and the base code that causes it (nothing > Mongrel in it at all). This code kind of simulates how Mongrel is > managing threads and locking Rails. Try out these graphs: http://pastie.caboo.se/10550 http://pastie.caboo.se/10551 On an otherwise quiet GNU/Linux system, I ran each of your scripts for half an hour in a ruby process into which I loaded hooks for malloc(), realloc(), free(), and brk() [1]. The red line is the amount of memory ruby is consuming from the heap. The green line is the amount of memory the heap implementation is consuming from the system. As Ara described in another message in this thread, Unix-y heap implementations (including glibc's) typically get new process memory from the kernel by calling brk() to extend the process's data segment. This can increase the apparent memory usage of code because the heap implementation can only move the end of the data segment backwards (releasing memory to the system) as far as the last page containing in-use memory. In the graphs I've created we can clearly see that in both cases ruby's use of memory from the heap drops back around the baseline with each iteration. In the 'sync.rb' case, the heap implementation's use of memory from the system tracks user code heap usage pretty closely. In the 'mutex.rb' case, however, we see two anomalies: (a) the heap implementation's use of system memory stays right near the maximum amount taken by ruby from the heap, and (b) the maximum memory used is around twice that of 'sync.rb'. Ara also mentioned in another message that Mutex is faster than Synchronize, and hypothesized that this was leading to a pathological interaction with the ruby garbage collector when creating as many threads as your example does. My data suggest that the observed "Mutex memory leak" /is/ caused by Mutex's relative speed, but that it is due to an interaction with the system heap implementation rather than with ruby's garbage collector. I haven't track down the exact factors at work, but I'd guess that 'mutex.rb' is (a) allocating new memory within the time window the heap implementation holds onto unused system memory, and (b) under some circumstances sufficiently fragments memory to force the heap implementation to allocate more system memory than required, eventually bringing the process under the gaze of the OOM killer. I hope this helps. [1] I used glibc's provided hooks for malloc/realloc/free and wrote a quick-and-dirty custom trampoline for brk(). Yes, I know about ltrace, but it only caught brk() as a system call, which slowed things down by a factor of 1000 (!). And yes Solaris DTrace would have made this easy, but I made my graphs while waiting for the 2.5Gb Solaris Express DVD to finish downloading. :-) I can provide my code if anyone wants it. -Marshall