From: "M. Edward (Ed) Borasky" Date: 2007-11-09T13:31:55+09:00 Subject: Re: Ruby/Fastcgi going into uninterruptible after random periods of time nate wrote: > Hello there - > > I manage some servers that run Ruby on Rails on top of Apache 2.0 and > FastCGI. Been running them for a little over a year now, and recently > upgraded OS/Ruby/FastCGI and have encountered an unexpected problem that > I've been having trouble trying to trace down. > > The servers have 8 cores, 8GB memory, and run about a dozen ruby > on rails applications, each application has a dedicated apache > instance with dedicated fastcgi. > > Original configuration(stable, despite some memory leaks): > Fedora Core 4 32-bit > Apache 2.0.54 > mod_fastcgi 2.4.2 > Ruby 1.8.4 > - a wide assortment of approx 25 gems and other ruby add-ons > (I can provide a list if needed) > > The applications are all in-house apps. I didn't write them, I > just manage the environment. In this configuration the servers > are rock solid stable, though most active application leaks > memory like crazy and is auto restarted once or twice a day. > > > New configuration > CentOS 4.5 32-bit > Apache 2.0.52 > mod_fastcgi 2.4.2 > Ruby 1.8.5 > - same wide assortment of approx 25 gems and other add-ons that > were recompiled against the above version of ruby. > > The problem is after random periods of time(3-20 hours) two of > the most busy ruby apps on the boxes(account for 80% of the > load) start going into uninterruptible sleep("D") for no > apparent reason. There is no I/O problem, the boxes run at > ~20% CPU, have real fast local disk subsystems, most of what > these apps do are talk to databases(on other systems). Processes in the "D" state for significant periods of time tell me that there *is* an I/O problem. Are these Ruby processes? Are other processes/kernel threads, for example, "kswapd" or "bdflush" also in a "D" state? If you haven't already done so, install the "sysstat" RPM. That will give you a marvelous tool called "iostat". Run "iostat -x -t" with ten-second samples and collect the logs in a file. When your system starts getting "D" state processes, look at the time stamp and compare it with the "iostat" logs. If you see one or more disks with a utilization of 100 percent, you're waiting for disk. Another thing to look at is those memory leaks. I don't remember what Fedora Core 4 used for a kernel, but CentOS 4.5 is a 2.6.9 kernel IIRC. The memory leaks may be starving the system of memory. "vmstat" will tell you that -- if you see lots of swap in and swap out, your memory leaks are making your system thrash. > But at the same time we ran 64-bit Ruby > in the same test environment for more than a month and it > lasted less than an hour in production before we had to > roll back(system performed a good 60-75% slower under > load than in 32-bit). > > I've worked on this a good 6-8 hours over the past week, > trying to troubleshoot it, searching the net for anyone > else that might of had the same problem with no luck. > > In the mean time I've been playing with mod_fcgid, got > it running but some of the basic 'smoke tests' that run > against our app fail on a particular part of the app with > mod_fcgid but not with mod_fastcgi. Haven't had much time > to look in-depth on it, Ruby just reports 'transaction > aborted', for some reason(even after raising timeouts in > mod_fcgid). > > I'm sure this is a lot to take in but I'm grasping at > straws here, hoping that maybe someone has some insight, > of all platforms I would of thought the newer one would > of been more solid than the older one. Well ... I wouldn't say grasping at straws so much as trying to many things at once to really isolate the problem. You've got way too many variables and factors at play here. If you've got one system / configuration that's stable, clone it and quit futzing with it for a week or so. :) Then make one change at a time, so you know what broke it! Incidentally, memory leaks are defects -- period. Either the infrastructure or the application code is broken. If it's your application code, get your developers busy fixing it. If it's open source code, bring it to the attention of the relevant community. > And before anyone suggest ditch mod_fastcgi, I've already > thought about it and we probably will at some point, > though, until now it's worked surprisingly well(I've > read about lots of complaints about setting it up, for > us it's a breeze). Worst case I can roll back to the > older OS and config in a couple of hours but I really > want to ditch that old distribution.. I think that's *exactly* what you need to do -- roll back to a stable configuration and introduce changes one at a time. > > I plan to post to the fastcgi mailing list as well, still > waiting to get subscribed to it .. > > thanks in advance for any insights.. > > nate > > > > > Good luck!