From: Dominik Bathon Date: 2005-08-01T05:00:49+09:00 Subject: Re: QUESTION: Good data structure for list within list with duplicates? On Sun, 31 Jul 2005 19:21:03 +0200, basi wrote: > Hello, > I have a dictionary within a dictionary, where the inner dictionary may > have duplicate keys: > > "x" => > "aa" => 2 > "aa" => 3 > "bb" => "boo" > .. > > "y" => > "aa" => 5 > "cc" => "see" > "cc" => "sea" > .. > ... > > The first level has about 750 entries. > > First, what is a good structure for this? I've looked at hash within > hash, but I'm not up to adopting some of the ways to handle duplicate > keys in hash -- unless there's one that even a rank newbie like myself > can use. You can use hashes within a hash, where the values of your inner hashes are arrays. Your example would then look like: { "x" => { "aa" => [2, 3], "bb" => ["boo"] }, "y" => { "aa" => [5], "cc" => ["see", "sea"] } } To ease the construction of this you should create your outer hash like that: h = Hash.new { |h,k| h[k] = Hash.new { |h2,k2| h2[k2] = [] } } If you have never used Hash::new with a block, then please read "ri Hash::new": -------------------------------------------------------------- Hash::new Hash.new => hash Hash.new(obj) => aHash Hash.new {|hash, key| block } => aHash ------------------------------------------------------------------------ Returns a new, empty hash. If this hash is subsequently accessed by a key that doesn't correspond to a hash entry, the value returned depends on the style of +new+ used to create the hash. In the first form, the access returns +nil+. If _obj_ is specified, this single object will be used for all _default values_. If a block is specified, it will be called with the hash object and the key, and should return the default value. It is the block's responsibility to store the value in the hash if required. h = Hash.new("Go Fish") h["a"] = 100 h["b"] = 200 h["a"] #=> 100 h["c"] #=> "Go Fish" # The following alters the single default object h["c"].upcase! #=> "GO FISH" h["d"] #=> "GO FISH" h.keys #=> ["a", "b"] # While this creates a new default object each time h = Hash.new { |hash, key| hash[key] = "Go Fish: #{key}" } h["c"] #=> "Go Fish: c" h["c"].upcase! #=> "GO FISH: C" h["d"] #=> "Go Fish: d" h.keys #=> ["c", "d"] Then you can build your example: irb(main):022:0> h = Hash.new { |h,k| h[k]=Hash.new { |h2,k2| h2[k2]=[] } } => {} irb(main):023:0> h["x"]["aa"] << 2 => [2] irb(main):024:0> h["x"]["aa"] << 3 => [2, 3] irb(main):025:0> h["x"]["bb"] << "boo" => ["boo"] irb(main):026:0> h["y"]["aa"] << 5 => [5] irb(main):027:0> h["y"]["cc"] << "see" => ["see"] irb(main):028:0> h["y"]["cc"] << "sea" => ["see", "sea"] irb(main):029:0> h => {"x"=>{"bb"=>["boo"], "aa"=>[2, 3]}, "y"=>{"cc"=>["see", "sea"], "aa"=>[5]}} > The key query I'd need from the list is finding if in the embedded > dictionary a value exists for a key, where the key may have duplicates, > for example, Is "see" a value of the key "cc" in "y"? irb(main):030:0> h["y"]["cc"].include?("see") => true irb(main):031:0> h["y"]["cc"].include?("sea") => true irb(main):032:0> h["y"]["cc"].include?("s") => false Hope that helps, Dominik