From: Nanyang Zhan Date: 2007-05-08T17:47:58+09:00 Subject: Re: separate Chinese and English! with Ruby Eden Li wrote: > for UTF-8 encoded strings. Ruby > will just treat the string as a string of 8-bit bytes and give you > back whatever byte you asked for. > > irb(main):001:0> s = "大智若愚" > => "\345\244\247\346\231\272\350\213\245\346\204\232" > irb(main):002:0> s[0] > => 229 > irb(main):003:0> s.length > => 12 in RoR console, I can see the string I put in: >> s = "大智若愚" => "大智若愚" >> s[0] => 229 >> s.length => 12 Zev Blut wrote: > If the goal is to separate the western languages from the Japanese > Kanji and Kana, then it appears to not be too bad when using a lib > like this: > > http://raa.ruby-lang.org/project/moji/ > > http://gimite.net/gimite/rubymess/moji.html Thanks, Zev. but my current problem is about Chinese. I am going to figure out a way to separate Chinese string from a string mix with other characters. What I mean other Characters are alphabets from English or/and other languages, like Ô, é, á... (may I call them western words?) This string may be containing no Chinese: "String without Chinese" ,I don't need to do anything about it, other than identify such strings. "中文 Western Words" #Chinese characters + space + western words. "中文・另一些中文 western words" #Chinese characters may be separated by punctuations, or/and space like: "中文 前有空格 western words" Almost all Chinese phrases are at the beginning of the strings. But some may contain numbers, like: "2007年的日记 diary of 2007" or some time English or alphabets are used as part of Chinese phrases,like: "BB日记 diary of my baby" Eden Li wrote: > Nooo! Those are the first BYTES of the UTF-8 encoding of the > punctuation that you listed. Finally, I know what those number are. Thanks. >so if you remove them from a givenstring, you're going to get back a poorly encoded UTF-8 string In fact, I wanted to use those number to test whether a character is Chinese or not (if 'character[0]' fit the range of [226, 228, 229, 230, 231, 233, 239], then it was likely to be a Chinese). (Now I know it may be wrong.) Then depend on this judgment, if this part of string ( string would be splited by space, divided into parts at the beginning) containing more X%, say 60%, of this kind of characters, then I would mark this parts as Chinese phrase, then take it out of string. I still want to use this strategy. but As you point out, [226, 228, 229, 230, 231, 233, 239] are not safe to identify Chinese, is there any other easy way to identify Chinese characters? > If you want to split on those separators, then why not do so > explicitly? > > # fill up c as you've done below >>> "asdf;asdfasdf".split(/#{c.join('|')}/) > => ["asdf", "asdfasdf"] I don't get it. what this code does? -- Posted via http://www.ruby-forum.com/.