From: alexander Date: 2007-02-09T21:01:20+09:00 Subject: Re: Problem with Base64 decoding whee! thank you! the three byte sequence you pointed out at the start of the file was the culprit. i just needed to [3..-1] that out of the way and everything works perfectly now... (crossing my fingers now that the app that�s producing those files doesn�t put illegal characters somewhere in the middle of the files, but that hasn�t happened yet.) according to the rfc this still seems like a bug to me. is there anywhere i should report that bug (if it is one)? thank you guys again for looking into this! really made my day that it�s solved now. kind regards, alexander Brian Candler wrote: > On Fri, Feb 09, 2007 at 06:20:28PM +0900, alexander wrote: > >>thank you for your tips! >> >>and indeed using write instead of puts i atleast got the filesize right. >>sadly everything else is still wrong. >> >>i think the problem is definitely in the decoding, but i don�t even know >>where to start there since the resulting files vary to a great degree. > > > Firstly, use hexdump -C on both the output files. > > If they both start with FF FE 00 78 00 61 01 FF FF FF FF > then you know that the headers are right and it's the base64-decoded bit > which is wrong. > > >>>>all_bytes = Base64.decode64(IO.read("test.rgb")) > > > BTW there's a built-in alternative: > > all_bytes = IO.read("test.rgb").unpack("m")[0] > > But on your test file they give the same results. > > >>>>File.new("test_ruby.gd","w").puts(bitmap) > > > If this is a Windows platform, use "wb" instead of "w". However you say that > now you're using write instead of puts, the files are the same size anyway. > > >>>>i�ve uploaded the test.rgb file i�m using to here: >>>> >>>>http://rss.fork.de/test.rgb if that�s even needed :) > > > I can see two issues with that file: > > (1) It has no line breaks, but I don't think that matters. > > (2) It starts with the three-byte sequence ef bb bf, which is a unicode > character according to my editor. > > Stripping this off gives a completely different answer to the base64 > decoding: > > irb(main):027:0> a=IO.read("test.rgb"); nil > => nil > irb(main):028:0> b=a.unpack("m")[0]; b.size > => 46560 > irb(main):029:0> c=a[3..-1].unpack("m")[0]; c.size > => 46560 > irb(main):030:0> b[0..5] > => "\304\000\000={u" > irb(main):031:0> c[0..5] > => "\000\365\355\326\000\342" > > and perhaps this second one is the answer you're looking for. > > If so, I would say that unpack("m") is badly broken. Either it should give > an exception when presented with characters outside of the base64 set, or it > should ignore them. According to RFC 2045 section 6.8, > > The encoded output stream must be represented in lines of no more > than 76 characters each. All line breaks or other characters not > found in Table 1 must be ignored by decoding software. In base64 > data, characters other than those in Table 1, line breaks, and other > white space probably indicate a transmission error, about which a > warning message or even a message rejection might be appropriate > under some circumstances. > > I would consider the unicode BOM as "white space", but in any case it must > either be ignored or cause a warning or error; it must not cause the data to > be decoded wrongly! > > BTW, I did the above test under ruby 1.8.4 (2005-12-24) [i486-linux] from > Ubuntu 6.06. It's possible that it has been fixed in a later version. > > HTH, > > Brian. >