From: sukhchander Date: 2006-08-13T05:30:17+09:00 Subject: Re: Regexp Parsing -- What's the right way? Hi Simon, That's pretty cool. I was looking for a utility similar to Java's StringTokenizer. You just pointed it out. Ruby has so many things built in. It's very comprehensive. For larger regexp I assume you prefer the split/tokenize method? I went with the Regexp approach because it occurred to me first. Thanks. [sukhchander] Simon Kr�ger wrote: > skelastic@gmail.com wrote: > > Greetings, > > > > I'm trying to parse the following line: > > > > "00608 P 135 001 LEC Tu 2-5P 210 WHEELER Information Tech and Soceity 3 > > LAGUERRE" > > > > i've constructed the following regexp: > > /(\d{5}).(\D\s\w{2,4}).(\d{1,4}\s\D{3}).(\D{1,4}\s\d+).*(\d{1,4}\s\D{1,9}).(\w+).*(\d?).*(\w{1,14}).*/ > > > > with a input file i've successfully produced the following output: > > control# 00608 ---- correct > > course#: P 135 ---- correct > > section#: 001 LEC ---- correct > > day-hour#: Tu 2 ---- missing '-5P > > room#: 3 LAGUERRE, ---- should be 210 WHEELER > > course-name#: M --- IT and Soceity > > credits#: --- should be 3 > > prof#: 5 --- should be LAGUERRE > > > > i'm a novice to ruby and regexp. i would like to know if i'm taking the > > right approach. > > i'll eventually nail it but any hints or suggestions would be useful. > > > > appreciate the help. > > I would go with split in this case: > > t = "00608 P 135 001 LEC Tu 2-5P 210 WHEELER Information Tech and Soceity 3 > LAGUERRE" > a = t.split > #strip from the beginning > control = a.shift > course = a.shift + ' ' + a.shift > section = a.shift + ' ' + a.shift > hour = a.shift + ' ' + a.shift > room = a.shift + ' ' + a.shift > #strip from behind > prof = a.pop > credits = a.pop > #the rest is the name > coursen = a.join(' ') > > puts "control: #{control}" > puts "course: #{course}" > puts "section: #{section}" > puts "hour: #{hour}" > puts "room: #{room}" > puts "coursen: #{coursen}" > puts "credits: #{credits}" > puts "prof: #{prof}" > > cheers > > Simon