From: Brian Candler Date: 2007-02-09T18:47:34+09:00 Subject: Re: Problem with Base64 decoding On Fri, Feb 09, 2007 at 06:20:28PM +0900, alexander wrote: > thank you for your tips! > > and indeed using write instead of puts i atleast got the filesize right. > sadly everything else is still wrong. > > i think the problem is definitely in the decoding, but i don�t even know > where to start there since the resulting files vary to a great degree. Firstly, use hexdump -C on both the output files. If they both start with FF FE 00 78 00 61 01 FF FF FF FF then you know that the headers are right and it's the base64-decoded bit which is wrong. > >> all_bytes = Base64.decode64(IO.read("test.rgb")) BTW there's a built-in alternative: all_bytes = IO.read("test.rgb").unpack("m")[0] But on your test file they give the same results. > >> File.new("test_ruby.gd","w").puts(bitmap) If this is a Windows platform, use "wb" instead of "w". However you say that now you're using write instead of puts, the files are the same size anyway. > >> i�ve uploaded the test.rgb file i�m using to here: > >> > >> http://rss.fork.de/test.rgb if that�s even needed :) I can see two issues with that file: (1) It has no line breaks, but I don't think that matters. (2) It starts with the three-byte sequence ef bb bf, which is a unicode character according to my editor. Stripping this off gives a completely different answer to the base64 decoding: irb(main):027:0> a=IO.read("test.rgb"); nil => nil irb(main):028:0> b=a.unpack("m")[0]; b.size => 46560 irb(main):029:0> c=a[3..-1].unpack("m")[0]; c.size => 46560 irb(main):030:0> b[0..5] => "\304\000\000={u" irb(main):031:0> c[0..5] => "\000\365\355\326\000\342" and perhaps this second one is the answer you're looking for. If so, I would say that unpack("m") is badly broken. Either it should give an exception when presented with characters outside of the base64 set, or it should ignore them. According to RFC 2045 section 6.8, The encoded output stream must be represented in lines of no more than 76 characters each. All line breaks or other characters not found in Table 1 must be ignored by decoding software. In base64 data, characters other than those in Table 1, line breaks, and other white space probably indicate a transmission error, about which a warning message or even a message rejection might be appropriate under some circumstances. I would consider the unicode BOM as "white space", but in any case it must either be ignored or cause a warning or error; it must not cause the data to be decoded wrongly! BTW, I did the above test under ruby 1.8.4 (2005-12-24) [i486-linux] from Ubuntu 6.06. It's possible that it has been fixed in a later version. HTH, Brian.