From: Arun Kumar Date: 2009-08-23T19:46:23+09:00 Subject: Re: Parsing pdf files --00151757374ade23370471cccf17 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: quoted-printable hello Alex, Thank you. But I would like to point out that its not very accurate in maintaining the layout. I already tried it out. you can copy a pdf file fro= m evince to gedit, you will get a better accuracy of layout. What escapes me is how to do it programatically :) cheers & regards, Arun On Sun, Aug 23, 2009 at 3:49 PM, Axel Etzold wrote: > Dear Arun, > > there is a command-line tool pdftotext, which you can use with encoding > specifications and also with a "-layout" option, which will preserve > line breaks. > The list of possible encodings > > pdftotext -listenc > > does not include iscii-1988, so probably, you'll be out of luck > if the original document is not in Unicode (maybe you can use iconv > on the result of pdftotext). > > I found a utf-8 encoded web page in Hindi, printed it to a pdf file, used > pdftotext on it, and opened it in the SciTE editor, specifying the > encoding as UTF-8. Most of the symbols are recognized correctly, > but some are not ...(vowels? combinations of letters?) > > I'm sending the screenshot as an attachment to your email address. > > Best regards, > > Axel > -- > Jetzt kostenlos herunterladen: Internet Explorer 8 und Mozilla Firefox 3 = - > sicherer, schneller und einfacher! http://portal.gmx.net/de/go/atbrowser > > --=20 || =E0=A4=B6=E0=A5=8D=E0=A4=B0=E0=A5=80 =E0=A4=9C=E0=A4=BE=E0=A4=A8=E0=A4= =95=E0=A5=80=E0=A4=B0=E0=A4=98=E0=A5=81=E0=A4=A8=E0=A4=BE=E0=A4=A5=E0=A5=8B= =E0=A4=B5=E0=A4=BF=E0=A4=9C=E0=A4=AF=E0=A4=A4=E0=A5=87 || --00151757374ade23370471cccf17--