From: Carlos Date: 2007-01-05T07:10:08+09:00 Subject: Re: ruby iconv UTF-8 to ISO-8859-2 (Polish) [Tobes , 2007-01-04 21.50 CET] [...] > Here's a test case > > On web page after being loaded from DB: "Wyślij" [This is correct!] > In MySQL Analyser: "WyÅ›lij" [bad, even though MySQL analyser is > UTF-8] > In Interactive Ruby (IRB) printed to console, after loading from DB: > "Wy┼ølij" [expected in a DOS prompt!] > In IRB unpacked, after loading from DB: [87, 121, 197, 155, 108, 105, > 106] > > So, I can see that the character "ś" must correspond to the 3rd and > 4th bytes of "Wyślij". > > Looking at the Ruby help, I see I can do this > > p str.unpack("U*") to get the UTF-8 characters, which gives: > > [87, 121, 347, 108, 105, 106] > > According to this, > http://www.fileformat.info/info/unicode/char/015b/index.htm, character > 347 is in fact a "ś". > > This would suggest that the database has UTF-8 text, and it's getting > into Ruby without corruption! Is this right? I think yes. > So, the question now is why doesn't Iconv convert my UTF-8 to Latin2 > correctly... That could just be because the original text can't be > converted due to additional characters outside of the Latin2 set. I still think the problem is not in the iconv library, but in the FPDF one. Think about it this way: "³" in UTF-8 ([179], byte sequence [194, 179]) translates to [179] in latin1. "ł" in UTF-8 ([322], byte sequence [197, 130]) translates to [179] in latin2. Now you provide a string with the byte 179 to the PDF library. How should it render that byte? As a "³" or as a "ł"? You must tell it how you want it by telling it which encoding you are using. If you don't tell it, it assumes you are using latin1; this is why you see a lot of "³" instead of "ł" in the PDF output at http://www.tobinharris.com/media/mtq38.jpg . The way to tell FPDF to use another encoding (for the PHP version, I suppose the Ruby one would be similar) is in the link I put in my previous message. Good luck. --