From: Axel Etzold Date: 2009-08-23T19:19:43+09:00 Subject: Re: Parsing pdf files Dear Arun, there is a command-line tool pdftotext, which you can use with encoding specifications and also with a "-layout" option, which will preserve line breaks. The list of possible encodings pdftotext -listenc does not include iscii-1988, so probably, you'll be out of luck if the original document is not in Unicode (maybe you can use iconv on the result of pdftotext). I found a utf-8 encoded web page in Hindi, printed it to a pdf file, used pdftotext on it, and opened it in the SciTE editor, specifying the encoding as UTF-8. Most of the symbols are recognized correctly, but some are not ...(vowels? combinations of letters?) I'm sending the screenshot as an attachment to your email address. Best regards, Axel -- Jetzt kostenlos herunterladen: Internet Explorer 8 und Mozilla Firefox 3 - sicherer, schneller und einfacher! http://portal.gmx.net/de/go/atbrowser