From: George Ogata Date: 2006-02-12T21:13:27+09:00 Subject: Re: OpenStruct problem "Timothy Goddard" writes: > I meant to say "The .hash (of the hash) depends on the order keys were > added...", but on second thoughts I think I was wrong. Two identical > hashes seem to have different .hash values even when the keys are added > in the same order. This obviously makes them completely unsuitable as > hash keys. Hash does not provide a #hash method at all; it falls back to Object#hash, which is essentially #object_id. But you'd be correct to infer that inserting keys in a hash, `h', in a different order could yield a different value for `h.to_a.hash'. You'll probably see the difference with: h1 = {} h2 = {} (1..1000).each{|i| h1[i] = true} (1..1000).sort_by{rand}.each{|i| h2[i] = true} p h1.to_a.hash == h2.to_a.hash p h1.to_a.sort.hash == h2.to_a.sort.hash The reason is that if two keys hit the same bucket, then which ever one was inserted last will appear at the head of the chain (ruby hashes use chaining). #each (and #to_a) simply scan each chain in turn and yields each element as it finds it. Thus insertion order can easily have an effect on the order of the array, and hence the array's #hash. As for usefulness, that will always depend on the application. Claiming it as "completely impractical for normal use" sounds akin to premature optimization to me, unless you're referring to a specific use case which you know well. You could also use a faster hash function which doesn't depend on the order, such as XORing values computed from each pair. I agree though, that a standard Hash#hash could be useful. It's come up before; I think the issue was how to handle the #default and #default_proc, especially the latter. I would probably favor using default_proc.object_id in the computed hash, even if it means two default_procs that always yield the same value end up giving different hashes.