From: Eric Mahurin Date: 2008-02-05T00:57:09+09:00 Subject: Re: Parsing JSON (#155) ------=_Part_7248_712601.1202140636600 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: 7bit Content-Disposition: inline On Feb 4, 2008 4:29 AM, Paolo Bonzini wrote: > On Feb 4, 11:01 am, Paolo Bonzini wrote: > > Here is my try using regexes. I use the "copy-on-write trick" from > > the suffix tree quiz: the regex is always anchored to the beginning of > > the string using \A, and the matched text is discarded using > > post_match. In some places where I don't want to discard I use (? > > =...). > > > > Using Eric's benchmark I get 36kb/sec, but I haven't benchmarked any > > other solution. > > > > http://pastie.caboo.se/147201 > > For what it's worth, I get 144kb/sec from tho_mica_l's solution, after > converting it to Ruby 1.8 like this: > > class JSONParser > > RXE = / > \[|\]| > \{|\}| > (:)| > (,\s*[}\]])| > ,| > ("(?>[^"\\]+|\\(?:u[0-9a-fA-F]{4}|[bfnrt"\/\\]))*")| > -?(?=\d)(?>0|[1-9]\d*)(?>\.\d+)?(?>[Ee][+-]?\d+)?(?=\D|$)| > true| > false| > (null)| > (?>[[:space:][:cntrl:]]+)| > ((?>.+)) > /xmu > > def parse(json) > ruby = json.gsub(RXE) do |t| > if !$5.nil?||!$2.nil? then invalid($5.nil? ? $2 : > $5) > elsif !$4.nil? then 'nil' > elsif !$1.nil? then '=>' > elsif !$3.nil? then $3.gsub(/#/, '\\\\#') > else > t > end > end > begin > return eval(ruby) > rescue Exception => e > invalid(json) > end > end > > def invalid(string) > raise RuntimeError, 'Invalid JSON: %s' % string > end > > end > Thank you. Now we can compare apples to apples with this on 1.8.6. I are the results with these two new benchmarks on my machine: ch/s author/gem ---- ---------- - Pawel Radecki (RE, mismatch) 3214 Justin Ethier (RE lexer + ruby eval, fixed number parsing) 4054 Eric Mahurin (Grammar0, no lexer, no parser generation) 4078 Eric I (Treetop, unicode broken) 6534 oksteev (Treetop, mismatches in benchmark) 8313 Clifford Heath (Treetop, had to remove handling of "\/") 54586 Eric Mahurin (Grammar, no lexer, v0.5) 137989 Paolo Bonzini (RE) 166041 Thomas Link (RE lexer + ruby eval, ruby 1.9 results) 220289 json 223486 Eric Mahurin (Grammar, no lexer, unreleased) 224823 fjson (uses C extensions) 333368 Thomas Link & Paolo Bonzini (RE + eval, unicode broken) 553081 Eric Mahurin (Grammar, no lexer, unreleased, w/ ruby2cext) 1522250 json (w/ C extensions) Looks like the above solution is fastest pure-ruby solution out there. It's interesting that the performance is better in 1.8.6 than 1.9. The "eval" solution is a nice trick as it uses ruby's C parser to act as the parser. But, this isn't viable for most other languages you might parse. Maybe fine for this JSON case though. The "eval" does seem a bit dangerous though. ------=_Part_7248_712601.1202140636600--