From: Robert Klemme Date: 2005-06-22T16:05:36+09:00 Subject: Re: Histogram-type Data "Florian Frank" schrieb im Newsbeitrag news:42B8D1A3.4040005@nixe.ping.de... > Charles L. Snyder wrote: > >> I have a text-parsing question - >> >>array1 = ["Joe lives in France", "Sue goes to Japan", "Hiro walked around >>France", "Bill moved to England"] >>a_countries = ["France", "Japan", "England", "Spain"] >> > puts a_countries.inject({}) {|h,c| h[c] = array1.inject(0) { |s,a| s + > a.scan(c).size }; h }.sort_by { |_,v| -v }.map { |a| "%10s%10u" % a } If I'm not mistaken your variant iterates array1 a_countries.size number of times. For large arrays this might be quite inefficient. This is what I'd do: counts = a_countries.inject({}){|h,c| h[c]=0; h} array1.each do |text| text.scan(/\w+/i) {|m| c = counts[m] and counts[m] = c+1} end You could even optimize away one hash lookups per found word by doing: class Counter def initialize(n=0) @n=n end def inc() @n+=1 end def to_int() @n end end counts = a_countries.inject({}){|h,c| h[c]=Counter.new; h} array1.each do |text| text.scan(/\w+/i) {|m| c = counts[m] and c.inc} end Kind regards robert