From: Sunny Hirai Date: 2006-11-27T14:58:48+09:00 Subject: Re: Two Advanced Ruby Performance Questions Hi All and thanks for the responses. Before I go on, I thought I should note that Vidar Hokstad is understanding my particular problem the best. Thank you for all the great responses but I would like to mention that I have a good grasp of how we will probably need to scale our application and the potential pitfalls. As I mentioned, we have already scaled one application to a good scale. I am looking for specific answers to the fill in the holes of my knowledge with Ruby. I am a Ruby newbie, but I have read three books on Ruby and a couple of them multiple times. I'm pretty much aware of most of what I want to know about the other parameters that you have asked me to consider. However, in no particular order, I'd like to address these issues to show you that I'm not looking for the "typical" answers as I mentioned in the OP. 1. This is a NEW application so many won't be switching. I will write the majority of it to start. It IS expensive for me to switch, but I'm thinking what each language will cost/save us 2-3 years down the road. Furthermore, I wrote all of the original code in the current application and will write all the core code for the new one. In other words, I will bear the brunt of the complexity which I'm currently determining whethere it is worht it. None of our developers except one actually knew ColdFusion from the start and were all taught ColdFusion on the job. All except two developers (not counting me) were hired less than a year ago in a 6 person team which means they also had to learn both a language and a framework. This is because it is difficult to find great ColdFusion developers. Rather, we hired smart people and taught them a new language. We will keep many developers with the old app and some will come over to the new app team. We are hiring two new people who will likely not know ColdFusion or Ruby anyways. I think ColdFusion is easier to learn but Ruby has a better syntax and I personally like its design philosophy better (though I like that of ColdFusion too). I am actually counting on, in the long term, a stable version of a VM, but am willing to wait a year to get it and just pack power against it in the shorter term. 2. As a database backed app, dropping to C is not needed for most of our application and having it split into two languages is costly from a development point of view so I'd like to keep the main app in one language only; however, there are performance oriented portions like a Photoshop like image processor that I've written in C# and .Net as a web service. It is Mono compatible but Mono seems to choke randomly. We used C# because of its strong underlying graphic framework, its Java like memory protection and still the ability to drop to UNSAFE mode where the custom filters, merging and effects code need to execute very quickly. We also have portions of interoperability in Java. Most likely we will also have a Lucene search engine in Java. All of these operate (or will operate) as web services. 3. I address the overhead of web services operations constantly. Even when working within the SAME language, I benchmark overhead and test the cost of each operation. For example, ColdFusion has a native WDDX (XML like) conversion format that we used to use to fold multiple fields into a single field for caching in a database. After testing, I wrote a custom encoding/decoding format that executes about 1-2 orders of magnitude faster if I remember correctly. I also weigh the cost of calling methods. In fact, one of the major reasons for wanting a switch to Ruby is that the object instantiation cost in ColdFusion is very high. It takes about 1ms per instantiation. This means that if I need to instantiate an object for say each row in a 50 row query for output, I have added a 50ms overhead to our project (in which we are aiming at under 100ms for total page execution times). And to cover the next response, YES, I know I shouldn't be doing this in an environment where object instantiation is expensive so I don't do this. Instead, I have three ways to instantiate an object which I've created in the framework, two ways which fake it in a manner that is 2 orders of magnitude faster; however, the syntax is ugly. It also adds to development time greatly because I need to decide every time which model I need to use and sometimes, the best model changes over time. I could use the fastest method (with the ugliest code) but it just introduces a layer of complexity and potential for bugs to the code which goes against my instincts. I constantly weigh app performance against developer performance. Most likely, however, is that I don't use OOP at all and inline the code but this results in bad reusability. Okay, now, that I've (hopefully) convinced you I'm attacking the problem at an atypical level, back to the regular program. ;) I found some of the information I wanted in not the eRuby or ERB pages but in mod_ruby. It suggests that ONE instance of Ruby executes to handle all the threads; however, it doesn't go into too much detail about how this is handled. So, for example, if I call require 'somelibrary' It is actually only included into the code once. However, it doesn't say anything about where scopes begin and end. The general feeling I'm getting is that there is very little documentation on the guts of Ruby and I'd like to learn about them without having to read the source code which I probably wouldn't understand anyways. For example, I'd like to know how multiple threads are handled. It appears that objects in the "global" scope are shared but objects in normal scopes are not. I found this in the FAQ for mod_ruby here: http://wiki.modruby.net/en/?FAQ#Why+are+changes+to+my+library+not+reflected+in+the+server%3F http://wiki.modruby.net/en/?FAQ#How+do+I+keep+an+object+instance+between+invocations+of+a+page%2C+for+example%2C+a+persistent+database+handle%3F But it is REALLY unclear how this all works. So for example, if I extend the original Array object, this does NOT persist between requests. But if I add the extension in a "require"d file, it then does persist but does NOT reload. I find this contradictory. It probably has something to do with this statement in the FAQ: --- You can't override classes in your mod_ruby scripts directly. (Instead, a new class will be defined.) Because mod_ruby scripts are loaded by Kernel#load(filename, true). If you have to override existing classes, please do it in a library, then require it from your mod_ruby scripts. --- But I'm not really sure how Kernal#load works underneath either. I'd also like to know how to deal with locking global variables for transactional use in a multi-threaded environment (e.g. web server) or if this is even possible. I still don't know whethere eRuby can be called from within Ruby or if it has to be called from the command line or through some sort of adapter. I feel like Ruby needs a "High Performance Ruby" book. There is one for MySQL and that is the only reason I had the confidence to make the decision to switch out of MS SQL Server. Knowing what I'm up against would help tremendously. Thanks for your feedback. If anybody knows anything more about the guts of mod_ruby and/or Ruby, please let me know. All the best, Sunny Hirai CEO, MeZine Inc. -- Posted via http://www.ruby-forum.com/.