From: Nuralanur@... Date: 2006-11-22T02:03:52+09:00 Subject: Text extraction from PDF files (non-European languages)...? -------------------------------1164128613 Content-Type: text/plain; charset="US-ASCII" Content-Transfer-Encoding: 7bit Dear all, is there a way of extracting text from a PDF, if the latter is in some non-European language, such as Arabic or Chinese? Under Linux, I have been able to use Ruby in conjunction with pdftotext for English and other Latin1 encoded texts - with some problems sometimes for special characters, but it doesn't seem to work for Unicode ... Is there a Ruby way to do this ? Thank you! Best regards, Axel -------------------------------1164128613--