From: Brian Candler Date: 2003-03-11T07:00:16+09:00 Subject: Re: DRB and threads On Mon, Mar 10, 2003 at 01:03:27PM +0900, ahoward wrote: > > Don't trap the SIGHUP. The child dies straight away, it gets respawned > > straight away by Apache (which will keep a minimum of one child per fcgi > > around unless you configure it otherwise), so the next request is handled > > successfully. > > the problem is that this drops the current request, if you are using a > transactional database that might be fine - otherwise... ... the client gets half a page, which IMO is no worse than getting a 500 Server Failed error :-) > > I mean using the drb protocol over HTTP as an API: e.g. front-end server > > talks to the world, and talks DRB-over-HTTP to the back end system, which > > has a pool of database processes run under fastcgi. It requires Ruby at both > > ends of course, but it should be a darned sight faster than SOAP or > > YAML/OKAY, and is *so* easy to use because you just make object calls on the > > front-end (which magically perform actions on the backend) > > i'm fuzzy on why you need fastcgi for the database backend? Well, it's like this. If I run a raw DRb server, I get into threading issues. Firstly, the way that a DRb server works, the object being served may get multiple method calls simultaneously from concurrent threads as multiple requests come in. So you can't just write: db = DBI.connect('dbi:foo:','bar','baz') DRb.start_service('druby://localhost:9000', db) because that would be highly dangerous. That can be fixed by wrapping it in an object which serialises the requests. In fact the that object can hold an array of separate DBI handles, which is even better: in principle, one client doing a long query then doesn't starve out other clients. In practice though, calls to DBI libraries are likely to block Ruby's threading. Maybe some don't, but at best it will depend on what type of DB backend you are using. That is the killer for me. And finally, on a multi-CPU box, I don't want it all taking place on one CPU. The solution then is to fork into multiple processes, each handling one request at a time. I didn't fancy writing the code to do that; furthermore, if doing it in Ruby I'd also have to modify DRb to work over a stdin/stdout pipe for the parent to pass messages to the child. Now, mod_fastcgi handles process creation automatically, even dynamically sizing the pool, and passing requests to individual children (using the fastcgi protocol). This seemed to be ideal. I made a small mod to DRb so that it does not accept() another incoming request until the current one has completed, and hey presto. My other solution would have been mod_ruby, which I've not tried yet - then I get one database handle for each Apache worker process, which also gives me the concurrency I need. > > Interestingly, the Ruby MOD_FCGI/CGI module itself isn't particularly > > speedly when compared with raw fcgi: > > there are **alot** more function calls in mod_fcgi... I think the overhead is in the Ruby standard CGI library, not in mod_fcgi. The attached program on my little laptop takes about 0.10 seconds per iteration. It turns out that the vast majority of this is the 'html3' bit, since it does a whole load of dynamic method additions: case type when "html3" extend Html3 element_init() extend HtmlExtension Changing CGI.new('html3') to CGI.new(nil) makes it run approximately 180 times faster, at 0.00055 seconds per iteration :-) The moral is: if you can avoid using CGI's tag-generating feature (which I don't think is friendly anyway) then you get a big improvement gain... Regards, Brian. require 'cgi' class Reader def initialize(str) @str = str end def read(n) res = @str @str = nil res end def binmode; end end ENV['REQUEST_METHOD']='GET' N = 20 src = (1..N).collect { Reader.new("foo=bar") } start = Time.now src.each do |s| $stdin = s CGI.new('html3') end puts "Per iteration: #{(Time.now - start)/N}"