From: "Jesús Gabriel y Galán" Date: 2010-08-02T17:52:56+09:00 Subject: Re: my script just read one line? On Sun, Aug 1, 2010 at 11:10 PM, Junhui Liao wrote: > Hi, Jesus. > > Thanks a lot for your help! > I modified a little to the script and  make it running as expected. > > Here is the code: > > def write_line_to_file every_line, base_time = Hash.new(0) >  every_line.each_slice(2).with_index do |(time,signal), index| >    File.open("header_split_#{index}"+".tsv" , "a") do |f| >      f << "#{time.to_f - base_time[index].to_f}\t#{signal}\n" >    end >  end > end > > #count = 0 >  first_line = file.readline.chomp.split("\t") >  # counter +=1 >  # if counter >= 2 >  # puts "here!" >  first_line_times = first_line.each_slice(2).map{|time,signal| time} >  file.each_line do |record| >    line_data = record.chomp.split("\t") >      write_line_to_file line_data, first_line_times >  end >  end > > > However, there existed two items need to be improved at least. > Item 1, this code took ~2 hours to save into 4096 files. > BTW, the original tsv file is around 250M. I wonder if there exist > some tricks to make it speed up? Maybe you can read it completely in memory, reorganize the contents per file, and then write each file at once. I think that should speed it up, although it implies a complete refactor of the code. > Item 2, the original data has 21 lines header. Although it could be > deleted then read by the script. But I do want to update the script > to make it exclude the fist 21 lines header. If you do a first file.readline after opening the file, you will read the first line. Then continue with what you already had. Jesus.