From: Bill Kelly Date: 2001-11-07T02:28:43+09:00 Subject: [ruby-talk:24448] Re: Joys of eval Hi Rik, From: "Rik Hemsley" > > The delimeter set is variable with each invocation, but never > '"', '(' or '\' (which are handled specially be the tokenizer > anyway) > > Yes, you can have \" within "" > > I'm actually down to only three calls in the whole parser, having > replaced everything else with regexps while porting to Ruby. > > tokenize(@strRep, ",\n\r", true, false) > tokenize(@strRep, " \n\r", false, true) > tokenize(@strRep, "@", false, false) > > Note: last two params mean 'skip comments' ( anything in () ) > and 'quoted tokens' (anything in "" should be treated > as a token, keep the surrounding quotes) > > e.g. > > Delimiter: . > Input: The.(quick)."brown.fox".jumps.("over").the."lazy.\"dog\"" > Output: <"brown.fox"> <"lazy."dog""> Here's an approach that constructs a tokenizer object given the delimiter string and skipComments/quotedTokens flags, which can then be used repeatedly to tokenize strings. Note, I haven't tried any timing tests with it. I'd be curious how much time the x.gsub! part adds to the process, for instance... Also, the last couple unit tests document its current behavior in a couple situations where I wasn't sure what the desired behaviour necessarily was. Regards, Bill ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ require 'runit/testcase' require 'runit/cui/testrunner' class Regexp def to_s; inspect[1..-2]; end # allow easy compositing of regexps end class Tokenizer ZapBackslashes = /\\(.)/ def initialize(delimStr, skipComments, quotedTokens) @scanner = construct_regexp(delimStr, skipComments, quotedTokens) end def tokenize(str) tokens = str.scan(@scanner).flatten.compact tokens.each {|x| x.gsub!(ZapBackslashes, '\1')} end def construct_regexp(delimStr, skipComments, quotedTokens) tokenCh = /(?:\\.|[^#{delimStr}])/ quotedTokenCh = /(?:\\.|[^"])/ commentTokenCh = /(?:\\.|[^)])/ comment = /\(#{commentTokenCh}*\)/ if quotedTokens token = /"#{quotedTokenCh}*"|#{tokenCh}+/ else token = /#{tokenCh}+/ end if skipComments /(?:#{comment}|(#{token}))/ else /(#{comment}|#{token})/ end end end class TestTokenizer < RUNIT::TestCase def testTokenizer # test stripping comments t = Tokenizer.new(".", true, true) a = t.tokenize 'The.(quick)."brown.fox".jumps.("over").the."lazy.\"dog\""' assert a == ['The', '"brown.fox"', 'jumps', 'the', '"lazy."dog""'] # test kept comments t = Tokenizer.new(".", false, true) a = t.tokenize 'The.(quick)."brown.fox".jumps.("over").the."lazy.\"dog\""' assert a == ['The', '(quick)', '"brown.fox"', 'jumps', '("over")', 'the', '"lazy."dog""'] # test escaping a = t.tokenize 'ignore.\"escaped.quotes\".here' assert a == ['ignore', '"escaped', 'quotes"', 'here'] a = t.tokenize 'ignore.\(escaped.comment\).here' assert a == ['ignore', '(escaped', 'comment)', 'here'] a = t.tokenize 'before."embedded \"escaped quote".after' assert a == ['before', '"embedded "escaped quote"', 'after'] a = t.tokenize 'before.(embedded \)escaped close-comment).after' assert a == ['before', '(embedded )escaped close-comment)', 'after'] # test backslash de-escaping a = t.tokenize %Q{hello\\"there\\\\this\\\\\\\\morning} assert a == [%Q{hello"there\\this\\\\morning}] # test multiple delimiter chars t = Tokenizer.new(" \\r\\n", false, true) a = t.tokenize %Q{ this \n(is\ra \n) \r "\n test "\r\n} assert a == ['this', "(is\ra \n)", %Q{"\n test "}] # does quote-handling need to happen within a comment? # this test assumes not: t = Tokenizer.new(".", false, true) a = t.tokenize 'ruby.("my)"dear)"' assert a == ['ruby', '("my)', '"dear)"'] # what's the expected behavior when quotedTokens==false? # here's the current behavior: t = Tokenizer.new(".", false, false) a = t.tokenize 'quotes."not.special".here' assert a == ['quotes', '"not', 'special"', 'here'] end end RUNIT::CUI::TestRunner.run(TestTokenizer.suite) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~