From: Mauricio Fernandez Date: 2006-01-12T10:51:39+09:00 Subject: Re: Re-post: Parsing Japanese Language and Some Ruby Trivia On Thu, Jan 12, 2006 at 10:23:11AM +0900, Michael Sullivan wrote: > Hi, > > I posted this last night and probably didn't hit the correct > audience. I got one relevant answer and need to go check the > packages recommended. But to those in Asian Time zones, I'll ask > again about parsing Japanese text. > > And to clarify, I am looking for a way to extract "words" from the > text for cataloging in a database. http://raa.ruby-lang.org/cache/ruby-chasen/ -> It looks old and the source code is not very enticing, though. This is probably better: http://chasen.org/~taku/software/mecab/bindings.html I wrote something similar to what you want long ago; for some reason I ended up parsing the output of mecab instead of using the bindings (can't remember why atm., nor if it still applies). The following (old, ugly) code collects some words (names, verbs, "i adjectives") from a utf8 string held in 'text': text = Iconv.iconv("eucjp", "utf-8", text).first temp = Tempfile.new "jphints" temp.puts text temp.close analysis = `mecab #{temp.path}` output = Iconv.iconv("utf-8", "eucjp", analysis).first temp.close! hints = [] output.each_line do |line| break if /\AEOS\s$/u.match line hint, nature, canonical = /\A(.+)\s+([^,]+),[^,]+,[^,]+,[^,]+,[^,]+,[^,]+,([^,]+),/u.match(line).captures # UGH case nature when %w[e5 90 8d e8 a9 9e].map{|x| x.to_i(16)}.pack("c*"), # name %w[e5 8b 95 e8 a9 9e].map{|x| x.to_i(16)}.pack("c*"), # verb %w[ e5 bd a2 e5 ae b9 e8 a9 9e].map{|x| x.to_i(16)}.pack("c*") # adj i puts "REG HINT #{hint} -> #{canonical}\t #{nature}" hints << canonical else # puts "IGNORED #{hint}" end end # now the words are in hints, as utf8 strings Hope this helps. -- Mauricio Fernandez