From: Jian Lin Date: 2009-05-10T11:29:30+09:00 Subject: Re: converting UTF-8 to entities like 剛 7stud -- wrote: > Jian Lin wrote: >> 7stud -- wrote: >>> 7stud -- wrote: >>> Whoops. With the 'm' switch: >>> >>> /..../m >>> >>> or >>> >>> /.{4}/m >> >> >> the problem is that some characters are converted to the correct >> 骼 etc, but some characters are not... you can try if you want... >> just go to Google News and get a China, taiwan, or hk news headline. > > Then why do you insist that you are trying to match any 4 bytes? no... the program converts the UTF-8 string into UTF-32, so that each character (code point) is 4 bytes long. And then the program process the end result, 4 bytes at a time, so that's why scanning 4 bytes at a time. -- Posted via http://www.ruby-forum.com/.