From: James Edward Gray II Date: 2007-03-30T09:58:16+09:00 Subject: Re: code too hashy, I think On Mar 29, 2007, at 4:44 PM, Giles Bowkett wrote: >> Any chance you could show us some trivial example data (just ten >> lines or so is fine and we only need the key fields) and how you want >> to access it. We might have better ideas when we see the >> specifics... > > OK, here's sample data and the script. I changed unique ID numbers and > made some subtle text changes as well to prevent the data from being > an NDA violation. > > http://gilesbowkett.com/blog_code_samples/muppet.csv > http://gilesbowkett.com/blog_code_samples/muppetimport.rb Usually we approach these problems with iterators. First I show some ways you might select subsets of data. This isn't as fast as Hash based access, but can be useful when you need to be able to view the data several different ways. If you still need the Hash indexes, I show how I would go about building those next. My hope is that something in here will give you some fresh ideas: #!/usr/bin/env ruby -w require "pp" $/ = "\r" # switch to the unusual line endings # read in the data Stuff = Struct.new(:family, :description, :color, :size_code, :sku) all_stuff = ARGF.inject(Array.new) do |rows, row| next rows unless row =~ /\S/ rows.push(Stuff.new(*row.split(/\s*,\s*/)[0..4])) end # find all by an sku when needed pp all_stuff.select { |s| s.sku.include? "35860466" } puts # find all entries of a certain size pp all_stuff.select { |s| s.size_code == "M" } puts # or group the data by size by_size = all_stuff.inject(Hash.new { |size, a| size[a] = [] }) do | grouped, s| grouped[s.size_code] << s grouped end pp by_size __END__ Hope that helps. James Edward Gray II