From: MonkeeSage Date: 2007-12-31T16:15:01+09:00 Subject: Re: using reg expr with array.index On Dec 30, 7:55 am, Esmail wrote: > MonkeeSage wrote: > > > No problem. What is really cool about the ruby community is that > > everyone is willing to help (and to learn!). I've never found a better > > programming community. So don't feel like any question is dumb or that > > you're asking too much. :) > > Thanks, I know that's the spirit of usenet, but there are some > groups where that idea has been lost unfortunately. I can hack > code together to get it to work in Ruby, but I want it to be a > good solution too :-) Ruby is sufficiently different from the > other languages I have experience with. For me, one of the greatest things about the ruby community is that the people who write books on the language, and even the language designer himself (matz), take the time to answer questions and interact with the community. Of course, you'll see a few RTFM replies now and then, in response to "do my CompSci homework for me" type posts; but on the whole, there really is no "ivory tower" in the ruby community. We're all just trying to learn and grow as programmers, and it's pretty much a level playing field. > > Without making assumptions about ordering and such, I'm not sure it's > > possible to avoid multiple iterations (and probably polynomial time) > > if you want to roll your own #uniq method to return an array of (or > > otherwise process) duplicate elements. Off the cuff, I'd say that > > something like this is probably the most efficient (but please correct > > if there's a better way): > > > def no_dups(path) > > seen = [] > > dups = [] > > IO.read(path).split(">").each { | item | > > if seen.include?(item) > > dups << item > > # or, for example... > > # puts %{Removed dup: >#{item.split("\n")[0]}} > > else > > seen << item > > end > > } > > [seen.join(">"), dups] > > end > > fixed, dups = no_dups("testfile") > > > Ps. I think google is indenting the ">" because it thinks it's the > > start of a quote. > > > Regards, > > Jordan > > This looks very much like what I wrote: GMTA, heh. ;) > ###################################### > # strip \n from data, find name of > # sequences and duplicate names > ###################################### > def processNames(data) > > names=[] > dups=[] > > data.each do |line| > > line.chomp! > > # find line with > > if line[0,1] =='>' > if !names.include?(line) > names.push(line) > else > dups.push(line) > end > end > > end #do > > return names, dups > end > ###################################### > > begin > if ARGV.length != 1 > puts "need to supply one command line arg" > else > file=File.open(ARGV[0]) > data=file.readlines > end > rescue > puts "Could not open file \"#{ARGV[0]}\"" > exit 1 > end > > #find names and duplicates > names, dups = processNames(data) > > I am using this code to process bioinformatics data in fasta format > (in case anyone's curious). I know there's a bioruby somewhere (I think) > but I am using this also as an opportunity to learn more Ruby. > > Esmail Ps. ruby will normally close open file handles on garbage collection or in finalization, but just in case of some catastrophic failure (what, I'm not sure), it's usually considered good practice to close file handles manually: ... file=File.open(ARGV[0]) data=file.readlines file.close ... Regards, Jordan