From: Simon Strandgaard Date: 2005-03-11T09:33:29+09:00 Subject: Re: Malformed UTF-8? On Fri, 11 Mar 2005 09:05:11 +0900, Ian Macdonald wrote: > One such allegedly bad string is the following: > > irb(main):001:0> "\032p\210\004n\306\271\310gY\002".unpack("U*") > ArgumentError: malformed UTF-8 character > from (irb):1:in `unpack' > from (irb):1 > > This is supposed to be Japanese. Can a Japanese reader please confirm > that this is, indeed, malformed UTF-8? I need to be sure that the bug > does not lie with Ruby before I get back to our calendar admin and tell > him to go and pester Oracle. the substring "\210\004" is invalid UTF8. in hex its [0x88, 0x04]. 0x88 has its uppermost bit set, so this is a dual byte sequence. 0x04 is not a valid continuation byte (upper bit should have been 1). -- Simon Strandgaard