From: Benjamin Thomas Date: 2009-10-23T07:42:39+09:00 Subject: Re: merging partial, incomplete data files James Edward Gray II wrote: > First, you could load one giant Hash (indexed by ref) of Arrays or > Hashes with the data as you read it, then just dump the Hash to disk. > This is the easiest, but it requires you to have all data in memory at > once. Are the files small enough that you can do that? Yes, I don't think that would be a problem. This is what I have so far : ############################################################################# def process_file(file) header = [] processed = [] opened = File.open(file, "r").readlines # store in array -- index=line opened.each {|line| processed << line.strip.split(";")} # remove separator and "\n" header = processed.shift return header, processed # the array is clean end def build_hash(header, arr, col_pos) hash = {} arr.size.times {|counter| hash["#{arr[counter][0]}_#{header[col_pos]}"]=arr[counter][col_pos] } return hash end hdr1, arr1 = process_file("file1") hdr2, arr2 = process_file("file2") p hdr1, arr1 puts p hdr2, arr2 hash1 = build_hash(hdr1, arr1, 0) hash2 = build_hash(hdr1, arr1, 1) hash3 = build_hash(hdr2, arr2, 0) hash4 = build_hash(hdr2, arr2, 1) puts p hash1, hash2 puts p hash3, hash4 puts hash = {} merge1 = hash1.merge(hash2) merge2 = hash3.merge(hash4) hash = merge1.merge(merge2) p hdr1 + hdr2 p hash ######################################################################### Outputs to : ["ref", "qty"] [["A", "10"], ["B", "20"], ["C", "30"], ["E", "5"]] ["ref", "price"] [["A", "100"], ["D", "150"], ["C", "200"], ["B", "75"]] {"A_ref"=>"A", "C_ref"=>"C", "B_ref"=>"B", "E_ref"=>"E"} {"A_qty"=>"10", "C_qty"=>"30", "B_qty"=>"20", "E_qty"=>"5"} {"A_ref"=>"A", "C_ref"=>"C", "B_ref"=>"B", "D_ref"=>"D"} {"B_price"=>"75", "A_price"=>"100", "D_price"=>"150", "C_price"=>"200"} ["ref", "qty", "ref", "price"] {"B_price"=>"75", "A_ref"=>"A", "A_price"=>"100", "C_ref"=>"C", "D_ref"=>"D", "B_ref"=>"B", "A_qty"=>"10", "D_price"=>"150", "C_price"=>"200", "C_qty"=>"30", "E_ref"=>"E", "B_qty"=>"20", "E_qty"=>"5"} ############################################################################ Robert Klemme wrote: > I'd say, you should look into classes Hash and Set. If you have both > files in a Hash each you can do... Thanks, I will look into "Set" which I don't know about. Gyepi SAM wrote: > However, I would strongly suggest that rather than doing the work yourself, > generate code to script sqlite or other database system. Yes, I have thought about that. I would guess it would be the most efficient solution in terms of ressources and speed. However I don't mind taking the time and making an effort to understand data manipulation. > Rather than using generating two arrays, I would read one file > into a hash, keyed by the common key, 'ref', then read the second file > and match up its record, if any, in in other file, and output the > combined Is what I've done similar to what you had in mind ? Thanks all for the input. -- Posted via http://www.ruby-forum.com/.