From: Phillip Hutchings Date: 2006-06-23T05:57:59+09:00 Subject: Re: ascii representation of unicode string? On 6/23/06, darren kirby wrote: > Hello all. > > I am unpacking some unicode strings from a binary file. I have a string like: > > "W\000M\000/\000T\000r\000a\000c\000k\000N\000u\000m\000b\000e\000r\000\000\000\003" > > and I need to turn it into: > > "WM/TrackNumber" > > When I 'puts' the string it prints fine but I need to assign it to a variable, > and when I try something like: > > require 'jcode' > $KCODE = 'UTF8' > s.each_char { |ch| print ch } > > it will print each char but return the original unicode string. And when I > try: > > n = "" > s.each_char { |ch| n += ch } > > The entire unicode char is being added to n. > > So can I extract an ascii representation of this string? I will admit I don't > know the first thing about unicode and I may be totally lost here... > > Thanks for consideration. That's not UTF-8, that's UTF-16 little endian without a BOM. If you know the string is pure ASCII, just UTF-16 encoded you can just do s.gsub(/\000/,''), but this will break any non-7bit characters. -- Phillip Hutchings http://www.sitharus.com/