From: John Joyce Date: 2007-05-09T01:20:58+09:00 Subject: Re: separate Chinese and English! with Ruby On May 9, 2007, at 12:22 AM, Nanyang Zhan wrote: > Michal Suchanek wrote: >> I guess this should give you what you want: >> >> irb(main):001:0> s = "大智若愚 asdfaf sdgs" >> => "\345\244\247\346\231\272\350\213\245\346\204\232 asdfaf sdgs" >> irb(main):002:0> s.unpack "U*" >> => [22823, 26234, 33509, 24858, 32, 97, 115, 100, 102, 97, 102, 32, >> 115, 100, 103, 115] > > Michal, Thanks! > Chinese character start from 4e00 to 9fa5 at the unicode table, and > CJK > symbols and punctuation range from 3000 to 303f. > > I just used my strategy combining this new way (unpack "U*") to > identify > Chinese, It picked out 100% Chinese phrases from the strings. (1000 > strings are tested) > > All of you that have replied and helped, thank you! Enjoy! NZ, could you share your final combined code? It might be useful to anyone using CJK, since Ruby originates in Japan that means a lot of people might find it useful. Might consider making a little gem out of it.