From: Robert Klemme Date: 2006-11-16T18:50:05+09:00 Subject: Re: *Fast* way to process large files line by line On 15.11.2006 23:35, devesh wrote: > Hi Everyone, Thanks for replying. > > First a couple of stupid things that I am doing, then a couple of more > questions. > > Stupid things: > i. I parse each line, but that essentially amounts to spliting it and > such. So maybe I should be better off using scanf kind of stuff and > also compiling regexps once and using them for each line. I heard that > helps. Using regexps inline is the most efficient way in Ruby. You can find numerous benchmark postings in this group. > ii. Going thru my hash H to delete keys is rather stupid. A better way > is to just dump it and trash it. I don't exactly understand what you mean but that's probably because I still do not have a clear picture of the processing your are doing. > iii. When I said it takes 8-9 hours, I meant 8-9 hours to read each > line and parse it (with my rather non efficient parse function). And I > did this one file at a time in sequence > > iv. I will definitely try to profile my code. I did use -rprofile but > that kind of revealed the obvious that most of my time was spent in > reading file loop > marshalling > parsing. Plus it was rather slow for > even 100MB of stuff. I don't think you have a CPU issue - this looks rather like an IO issue. > Oh and btw, the disk I am using is an LVM mapped ext3 local disk. The > server runs Centos. And I think it is pretty well managed. Sounds good. > Now my questions: > > i. I was earlier doing this: Read each file in sequence, and then > write out a similar hash, containing data from one src for each time. > Even with this hash, I used to dump it to temporary files (one per hash > key = time instant) and then after doing this for every file (src's). I > would then merge all this data in the temporary files into what I want. > This was slow, (but in retrospect not as slow as what I am doing now by > reading in parrallel). The reason I though opening multiple files and > reading from them line by line would be faster is cos I wouldn't have > to open/close all those thousands of temporary files. Wait, did you say "thousands of temporary files"? Wow, *that* is likely going to kill performance. > I still don't see the reason why doing things in parrallel would screw > things up ? As my assumption is that the disk head is anyway pretty > wild as the system will involve a lot of IO b/w context switches. That entirely depends on the other processing going on on that machine. And there is no point in making it worse by accessing multiple files at the same time. > ii. Also why do you think that writing another file in the midst of > reading (one or more) one is a bad idea ?. The only way I can avoid > this is to chunk up files into smaller units and then process them, > writing their results onto temp files and then proceeding along further > with other files. Which one do you think is a better thing to do ? Writing is usually slower than reading. But it entirely depends on what you actually do. > I will now try the following (based on your suggestions): > i. Avoid heavy weight regexps in parsing lines Not necessarily a good idea - especially it seems your task is IO bound. You can easily verify with "vmstat 5" while you are doing your processing. > ii. Avoid marshaling/dumping This in itself is not slow. But if you do this into multiple small files you'll have problems - but not because of Marshal being slow but because of the IO. > iii. Focus on doing things a file at a time. > iv. If all else fails, I will try using something like sharkfish etc. > > Btw will using something like an mmap extension for ruby speed things > up for me ?. I don't believe so because your problem seems to be the IO access pattern. > Thanks a lot to all of you. I am sorry if I sound lame, but I am pretty > new to using ruby. Can you describe in more detail what you are actually doing? Kind of: are you combining data from multiple files? How long do you need temp data? What do you output - things like that. It may actually turn out that Ed's advice to use a RDBMS is the best solution because that gives you efficient IO. But to confirm or reject that I would have to have a better understanding what business problem you are trying to solve. Cheers robert