From: Alexander Lamb Date: 2005-10-21T22:47:28+09:00 Subject: Re: Threads, timing and HTTP --Apple-Mail-2--876048802 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=ISO-8859-1; delsp=yes; format=flowed Well, as you said: "good enough" is ok since I need to test a real =20 life situation. For example I have several apps calling some web =20 services on some given servers. Obviously if many apps call at the =20 same time the timing will be different, as is my probe in Ruby. =20 However, I had the feeling that by starting my two or three threads =20 exactly at the same time to ping some servers, the timing result is =20 not really correct. Using processes, of course several processes will =20= fight for resources but I am more in a real life situation. However, reading your replies, I thing a good approximation is to =20 slightly offset by 3-4 seconds each ping (e.g. not starting all the =20 threads at the same time). Then I have a very high probability, even =20 after several hours of pings, to have only one ping thread running at =20= a time thus giving me a good approximation of the time taken. Slightly off topic: what I am trying to do is to monitor the way our =20 systems work. We are very distributed and need to setup alarms if =20 some service goes down. A little bit like products such as BigBrother =20= but more application oriented. I saw on the agenda of Euruku05 : Using Ruby to monitor enterprise software from Sven C. Koehler. There are no slides or description but it could be something similar =20 to what I am trying to do (actually that we already did in a previous =20= version in Java but I wanted to simplify it and make it more =20 customizable). Can someone give me pointers or maybe even Mr. Koehler =20= if he is on this list? Thanks, -- Alexander Lamb Service d'Informatique M=E9dicale H=F4pitaux Universitaires de Gen=E8ve Alexander.J.Lamb@sim.hcuge.ch +41 22 372 88 62 +41 79 420 79 73 On Oct 21, 2005, at 3:26 PM, Robert Klemme wrote: > Robert Klemme wrote: > >> Alexander Lamb wrote: >> >>> On Oct 21, 2005, at 3:01 PM, Robert Klemme wrote: >>> >>> >>>> Alexander Lamb wrote: >>>> >>>> >>>>> Hello list, >>>>> >>>>> I am implementing a very simple script to ping web servers or >>>>> services (to monitor how our environment is functionning). >>>>> >>>>> Some production url's run on more than one host. Therefore, I =20 >>>>> start >>>>> a thread for each separate url. >>>>> >>>>> The function run in my threads is: >>>>> >>>>> def doPing(uri_string, probe) >>>>> s =3D uri_string >>>>> while true >>>>> begin >>>>> timeout(@seconds_before_timeout) do |timeout_length| >>>>> start =3D Time.new >>>>> begin >>>>> open(s) do |result| >>>>> if result.status[0] !=3D "200" >>>>> probe.addToLogFile([s,'ERR',0,result.status=20 >>>>> [1]]) >>>>> else >>>>> probe.addToLogFile([s,'OK',Time.new - =20 >>>>> start,'']) >>>>> end >>>>> end >>>>> rescue Exception >>>>> probe.addToLogFile([s,'ERR',0,$!]) >>>>> end >>>>> end >>>>> rescue Timeout::Error >>>>> probe.addToLogFile([s,'ERR',0,'timeout']) >>>>> end >>>>> sleep(@seconds_between_ping) >>>>> end >>>>> end >>>>> >>>>> However, this is a problem. Indeed, I want also to measure the =20 >>>>> time >>>>> (round-trip) it takes for the ping (these are only simple pings =20= >>>>> for >>>>> the time being). As you can see I get the local time before and >>>>> after the call. But this doesn't work with threads. Indeed, since >>>>> the process is shared among threads, the time will be dependent on >>>>> the number of threads I am running and not a correct view of the >>>>> actual time it takes to ping. >>>>> >>>>> I can't define the piece of code between the two times as critical >>>>> and only for one thread because if the open-uri blocks, it will >>>>> prevent another thread to ping another url in the mean time. >>>>> >>>>> Any idea? Maybe use processes instead of threads? >>>>> >>>>> >>>> >>>> Maybe you can exploit one of the result headers. Chances are that >>>> there >>>> is a timestamp somewhere. Then you *only* need to synchronize >>>> clocks on >>>> your machine and on servers... >>>> >>>> Or you switch to a single thread solution. I don't know your ping >>>> interval but if you don't need to ping too often and don't have too >>>> many >>>> servers that should be ok. You can create a simple scheduling that >>>> always >>>> picks the URL with the closest ping point... >>>> >>>> >>>> >>> I could go single thread (since indeed I am doing a ping per 30 >>> seconds more or less). However, if the first one I try hangs =20 >>> (until a >>> timeout for example), I am pushing back the time at which I will =20 >>> ping >>> the second url. Logically I would need to do something like "ping >>> each url one after another unless one of them seems to take longer >>> and then detach a thread to wait for the answer". >>> For the time being I will test forking a process. >>> >> >> You could also have a controller thread that watches your single >> testing thread. If the testing takes longer than n seconds (where n >> << timeout) it sets a flag for the current testing thread (with a >> thread local variable for example) and starts a new tester thread. >> > > Yet another idea: you make the testing semi critical. When a thread > starts testing it stores a timestamp somewhere. Every other thread =20= > checks > whether the timestamp is set and is only max n seconds away. If it's > longer, replace the timestamp with it's own timestamp and go =20 > ahead. If > we're still in the n seconds range, go on sleeping. > > robert > > > --Apple-Mail-2--876048802--