From: Robert Klemme Date: 2008-01-02T20:29:57+09:00 Subject: Re: using reg expr with array.index On 02.01.2008 05:52, Esmail wrote: > Robert Klemme wrote: > Hi there Robert, > >> Esmail (btw, is that a real name?), > > Yes, it is :-) .. there are a lot of different alternate spellings > of this, Ismael, Ismail are quite common. In the US I have also > seen (and unfortunately heard) Ishmael (as written in Moby Dick I > believe). Unfortunately, that pronunciation sounds to me as if > someone is saying "Shteve" instead of "Steve" :-) Ah, ok. I knew "Ismael" but did not know that it could also be spelled with an "E" in front. Thanks for the explanation! Learn something new every day... > Ok, back to Ruby ... > >> since you are dealing with files I would like to revisit this. Since >> files a) are slower to read and b) are potentially large - especially >> larger than main memory - you might want to look for different >> solutions. Basically, unless you need the file's contents otherwise >> you should try to avoid having to store the whole file in memory at >> one point in time and strive to process the whole file as seldom as >> possible (ideally only once). > > Agreed, file io is a bottleneck and it may not even be possible > sometimes to store a huge file in RAM. In this case I'm working with > the knowledge (? .. or pretty good assumption) that the files will fit > into memory, so the ease of implementation becomes a factor. Absolutely. > But I am always looking for alternative ways of doing this, just in > case I run into trouble with one way of solving/tackling a problem. > > The files contain multiple DNA sequences. Each sequence starts on a > line with a '>' in column 1 and its header, and then is followed by an > unknown number > of lines with data. If there is another record, it will start with > a '>' in column one etc. > > The problem is that since I am concatenating a number of different > sequences > into a large file, the possibility of duplicate sequences exists, which I > need to identify and eliminate. > > I could scan the file once to determine which sequences are duplicates > and then process the file a 2nd time eliminating those. In fact that was > my first approach, but then Jordan's suggestions were so much cleaner > and simpler that I went with them. My file aren't very large but it's > good to have some other approaches in mind. If the volume of data is always so that it will fit into memory (and if you think about it, it *has* to be in memory for the duplicate detection). Then I'd probably write a class (or find a class somewhere, probably in RAA) that will represent a sequence and have proper comparison methods (#==, #eql?, #hash etc.). That way you can even use a Hash for fast duplicate checks. And you might even be able to internally represent those sequences with less memory (compressed, encoded or whatever suits you best). For example, since you just need two bits to represent one element of the sequence you can achieve compression of factor 4 easily by not using a char per entry but just two bits. Kind regards robert