From: "yermej@..." Date: 2007-09-16T13:00:04+09:00 Subject: Re: XML parser; maybe ruby is too slow? On Sep 15, 7:49 am, nutsmuggler wrote: > Hello folks. > I managed to write a SGML parser with the hpricot library. As I > explained in a previous thread, I just need to compare source and > traget tags of translation memory files from IBM Translation manager. > The script now runs effectively, but I realised that it cannot cope > with large files; I tried to process TM file larger than 1MB and the > script took ages to generate the output. Should I switch to a compiled > language for this specific task? > At any rate, here is the script, it's very basic; please let me know > if I did something wrong or if its slowness is a necessary drawback of > ruby being interpreted. Cheers, > Davide > > #!/usr/local/bin/ruby > require 'rubygems' > require 'hpricot' > > $pattern = "server" > result = File.new("result.html", "w") > $stdout = result > puts " 'http://www.w3.org/TR/html4/strict.dtd'>\n > \n > \n > Ricerca di '#{$pattern}'\n > > \n > \n" > # per aprire lo stdin > # doc = Hpricot.XML(STDIN) > > doc = Hpricot.XML(open("bch01aad006_MEMORIA.EXP")) > doc.search("Source").each do |item| > if item.innerHTML =~ /#{$pattern}/ > highlightedSource = item.innerHTML.gsub(/#{$pattern}/, " class='pattern'>#{$pattern}") > puts "

EN: #{highlightedSource}

\n" > puts "

IT: #{item.next_sibling.html}

\n >
" > end > end > puts "" I haven't done any comparison testing, but if your *.EXP files are truly XML, Ruby libxml might be a better choice as it's just a wrapper around the libxml2 library (see http://libxml.rubyforge.org/). Jeremy