From: Robert Klemme Date: 2007-09-18T22:54:53+09:00 Subject: Re: finding string matches, in order, in a file 2007/9/18, Peter Bailey : > William James wrote: > > Peter Bailey wrote: > >> Hi, > >> I've got files I want to parse. I'm using a string scan routine that > >> populates an array. I need to pull the entries of that array out, in > >> order, eventually. I'm getting an array all right, but, I don't > >> understand its order. The first instance in the string, meaning the > >> whole file, is way down the list in the array. The first entry in the > >> array is an entry that's 300 lines deep into the file. Why isn't the > >> first instance in the string, the file, the first entry in the array? > > > > The file may have multiple copies of some entries, and > > your regexp may be botched. > > > >> > >> xmlfile.scan(/\n(.*)<\/issue>\n?/) > > > > I don't like the looks of that regular expression. Try this one. > > > > /\n(.*?)<\/issue>\n?/m > > > Thanks, William. I tried your regex, but, I'm still getting the first > entry as one that's 300 lines deep into the file. In fact, the results > look exactly the same to me. Still William's regexp is significantly better than the original one. You seem to be processing XML files. It may be that there is some white space between and that you are not prepared for. You can handle that by replacing \n with \s*. A completely different approach is to use REXML or another XML tool and use XPath search. This is way less error prone - but usually also slower. If you just want to extract these codes then a SAX parser approach might still be pretty fast. Kind regards robert