From: Dominik Bathon Date: 2006-04-11T07:01:18+09:00 Subject: Re: [QUIZ] Markov Chains (#74) On Mon, 10 Apr 2006 21:20:07 +0200, Ross Bamford wrote: > On Tue, 2006-04-11 at 02:15 +0900, Dominik Bathon wrote: > >> So, the morale of all this: >> - don't use symbols if they have to be converted to a string often >> - hash lookups might be slower than you think >> - premature optimization... > > I used symbols in my solution, but primarily because I was trying to > make the hash itself require as little memory as possible to allow for > vast bodies of input (which I admit I've not really tried yet) and I > figured a hash on symbols would be smaller than a hash on strings. Tha hash itself will have the same size, because it only stores VALUEs (which are just longs). But each of the string VALUEs will "point" to another "object" on the heap, while the symbol VALUEs don't have an extra "object" attached. > I guess this means that GC is pushed by lots of string objects created > and destroyed during the generation run, but I tended to aim for memory > efficiency over speed for this one, as long as it was running 'fast > enough'. Yes, memory efficiency was another reason for me to first try the "hash tree of symbols". But on the other side: the string objects are created anyway (before they are converted to symbols). They can be collected after conversion to symbol, but when you generate the sentence you are again creating many new string objects (Symbol#to_str generates a new string object every time, ruby stores only c-strings for the symbols), which wouldn't be necessary, if you had kept the strings in the first place. So, yes, symbols are more memory efficient for storing big frequency hashes, but they are slower for generating, so it's a tradeoff. By the way here are some numbers: order first final 2 7.380s 1.973s 4 6.279s 2.002s 6 8.031s 1.972s Those runs are for a 700K text file and 1000 sentences generated. I didn't measure the memory. Dominik