From: eden li Date: 2007-05-08T16:15:51+09:00 Subject: Re: separate Chinese and English! with Ruby Nooo! Those are the first BYTES of the UTF-8 encoding of thepunctuation that you listed. MANY Unicode characters (when encoded inUTF-8) can start with those bytes, so if you remove them from a givenstring, you're going to get back a poorly encoded UTF-8 string whichwill is definitely not what you want. If you want to split on those separators, then why not do soexplicitly? # fill up c as you've done below>> "asdf;asdfasdf".split(/#{c.join('|')}/)=> ["asdf", "asdfasdf"] On May 8, 2:54 pm, Nanyang Zhan wrote:>>> c=%w{“ ”。 , ! < { ; ‘ ! @ # $ % … * ( ) 一 俿 倀 凿 勿 叿 哿 囿 姿 寿 崁 忄忿 恘 扉 掵 曆 桶 檗 泗 濗 瀖 燿 狧 珗 痿 眀 秊 竗 篿 紀 翹 退 釽 鎷 閈 阀 韗 饧 骠 鶆 龥}>> => ["“", "”。", ",", "!", "<", "{", ";", "‘", "!", "@", "#", "$", "%",> "…", "*", "(", ")", "一", "俿", "倀", "凿", "勿", "叿", "哿", "囿", "姿", " 寿",> "崁", "忄忿", "恘", "扉", "掵", "曆", "桶", "檗", "泗", "濗", "瀖", "燿", "狧", "珗",> "痿", "眀", "秊", "竗", "篿", "紀", "翹", "退", "釽", "鎷", "閈", "阀", "韗", "饧",> "骠", "鶆", "龥"]>>> c.collect.map{|o| o[0]}.sort.uniq> => [226, 228, 229, 230, 231, 233, 239]>> maybe 226 to 239 is the range I need.