From: Tom Sawyer Date: 2002-07-04T15:08:08+09:00 Subject: Re: the pains of parsing thanks sean! your code turned out to be right on the money. only a couple of ever so minor adjustments and it worked like a charm. very elegent solution. i can't tell you how impressed i am. by the way, i am really happy to hear about the attributes in rexml! it means my xml:proof schema api is functional as is. i had been holding out ;-) so that's about ready to go. i'm not able to work on it presently, i'm busy with some other stuff (which you just helped me with!) but i'll get back to it soon and let you know if i run into anything. later and thanks again, tom On Wed, 2002-07-03 at 21:54, Sean Russell wrote: > Tom Sawyer wrote: > > > i find myself spending alot of time writing routines to parse complex > > strings. here's a mock example string of my current problem: > > > > a[c]{d}"e"f {{g}} [[h]]*i**j*"k" > > Hi Tom, > > This isn't the most beautiful or efficient code I've ever written, but the > method is only 22 lines long. You could probably get it down to 18 lines > if you obfuscated it a little more. Sorry for the perlisms. Once you > start using them, they're a hard habit to kick. > > You can't solve your problem with regexps alone. Oh, BTW, I wasn't sure if > you really meant you wanted whitespace to be stripped; if so, change the > line: > > string = $' > > to read > > string = $'.strip > > > def tokenizer string > tokens = { ?< => '>', ?[=>']', ?"=>'"', ?{=>'}', ?*=>'*', ?'=>"'"} > items = [] > while string.size > 0 > if tokens.keys.include? string[0] > end_index = string.index( tokens[string[0]], 1 ) > item = string[0..end_index] > items << item > string = string[end_index+1..-1] > while item.count( item[0,1] ) > item.count( tokens[item[0]] ) > end_index = string.index( tokens[item[0]] ) > item << string[0..end_index] > string = string[end_index+1..-1] > end > else > string =~ /(.*?)(?=[<[{"*'])/ > items << $1 > string = $' > end > end > items > end > > puts tokenizer( %Q{a[c]{d}"e"f {{g}} [[h]]*i**j*"k"} ).inspect > > > > ser@ender ~% ruby tokenizer.rb > ["a", "", "[c]", "{d}", "\"e\"", "f ", "{{g}}", " ", "[[h]]", "*i*", > "*j*", "\"k\""] > > > > -- > |.. "A scientist is one who finds interest in the kinetic energy of > <|> Jell-O moving at ridiculous velocities...an engineer is one who can > /|\ find a real-life application for such silliness." > /| -- anon > | >