From: Kirk Haines Date: 2004-04-26T23:22:13+09:00 Subject: Re: How do I scale large Ruby web applications? On Mon, 26 Apr 2004, Sascha Ebach wrote: > In my understanding the biggest problem in scaling Ruby (cpu wise) is > that it doesn't have native thread support, yet. What this means in > terms of a web application is that if you only have let's say 30 > concurrent users on a fairly new piece of Hardware this is not a > problem. But what happens if your site suddenly gets very popular and > you jump from 20 to 200 or even 2000 concurrent users? How do you scale > such a web app? If you were to program this web app in Java or any other > language which supports native threads you could simply throw more cpus > and ram at it. I am thinking of a blade server here. The more users you > get you simply stick another blade in your server and have your piece of > mind. As I understand it you cannot do something like that with Ruby. > Enter Distributed Ruby (DRb). I've been using Ruby for all of my web application and simple dynamic web page work for two years now, using Iowa as my framework. Here's my experience so far: I have yet to hit a performance bottleneck. However, my hardest hit "application" is actually a very large highly dynamic web site with several areas, both the for site users and the site owners, that resemble applications. Over the last two weeks it has served about 500000 dynamic pages. That level of load isn't even noticable on the server it is running on (which is FAR from a cutting edge machine). A large percentage of these hits involved MySQL database accesses, as well. On load testing that I have done on a development box, Iowa running on Ruby 1.8.1 on a 800Mhz Pentium III running Gentoo linux and going through Apache 2 currently seems bottlenecked by the way it handles communications between the web server and the Iowa application process. For a simple dynamic web page (< 1k in size), the current model on that hardware topped out at around 40 pages/second. For huge pages (40k - 50k) it was 10/15 pages/second. Both assumed that the actual dynamic processing load was not high. No lengthy calculations or querying hundreds or thousands of rows from a database and compiling them into a report. Nonetheless, I have been thinking both about performance and fault tolerance for a couple of years now (one sector I do a lot of work in is the financial sector, and they _really_ don't like seeing anything unavailable). My phase 1 expansion plan is to simply add an FCGI-like ability to run multiple application processes at once (which, via even Ruby's green threads, handle multiple requests at the same time quite well), tieing the process that is servicing the session to the session id. Once that is working well, I want to add a higher level proxy layer that will do the same thing except it will route requests to different machines, letting me scale to multiple servers and also letting me cope if a server goes down, at the cost of only losing whatever sessions happened to be on that server. Given the performance that I get right now, and the scalability that I expect to gain from a few small changes to my process communications and then from the ability to run multiple application processes, I don't xpect any scalability problems in my foreseeable future. Kirk Haines