From: James Edward Gray II Date: 2005-06-22T09:56:12+09:00 Subject: Re: Histogram-type Data On Jun 21, 2005, at 5:50 PM, Charles L. Snyder wrote: > Hi, > I have a text-parsing question - > > array1 = ["Joe lives in France", "Sue goes to Japan", "Hiro walked > around > France", "Bill moved to England"] > a_countries = ["France", "Japan", "England", "Spain"] > > desired result is unique list (hash table) of all countries > mentioned in > the sentences in array1, with frequency of occurence - eg, > > France 2 > Japan 1 > England 1 > Spain 0 See if this gives you some fresh ideas: irb(main):001:0> array1 = ["Joe lives in France", "Sue goes to Japan", "Hiro walked around France", "Bill moved to England"] => ["Joe lives in France", "Sue goes to Japan", "Hiro walked around France", "Bill moved to England"] irb(main):002:0> a_countries = ["France", "Japan", "England", "Spain"] => ["France", "Japan", "England", "Spain"] irb(main):003:0> counts = Hash.new => {} irb(main):004:0> a_countries.each { |c| counts[c] = array1.find_all { |line| line =~ /\b#{c}\b/ }.size } => ["France", "Japan", "England", "Spain"] irb(main):005:0> counts => {"Spain"=>0, "France"=>2, "Japan"=>1, "England"=>1} Hope that helps. James Edward Gray II