From: Robert Klemme Date: 2003-09-18T17:53:26+09:00 Subject: Re: Why IO#readlines does'nt accept a Regexp? "gabriele renzi" schrieb im Newsbeitrag news:uroimv475apt4pn9bqqtp2aa1s7ul4ngui@4ax.com... > as in the subject, I just noticed that readlines just accepts a string > as line Separator, and I wonder why it works this way. > Some explanations? Just a guess: normally it's not necessary and another reason might be performance, since the overhead of a regexp might be significant for large files. However, you can simulate it if you read a complete file into a string and then split with a regexp. > BTW, if I want to read a file in a array of 'words' I have to do : > > > File.new('myfile').gets(nil).split > > no better way ? For large Files this is more efficient: words=[] IO.foreach("myfile") do |line| words.push( *line.scan( /\w+/oi ) ) end If you have many repeating words you can save even more mem: cache = Hash.new {|h,k| h[k]=k} words = [] IO.foreach("myfile") do |line| words.push( *( line.scan( /\w+/oi ).map {|w| cache[w]} ) ) end > on a sidenote, what are the efficiency issue related to the use of > IO#each vs IO#foreach(anIO) vs a simple 'while line=gets..' ? Try ruby -profile with each method and see what happens. I'd guess that there is not much difference. Regards robert