From: Phillip Gawlowski Date: 2011-05-18T06:45:38+09:00 Subject: Re: Pdf Parsing Challenge On Tue, May 17, 2011 at 11:38 PM, Felipe Espinoza wrote: > I need to do this automatically, I'll be doing it for a lot of papers > and then take that data to a database Unless the papers are all (near) identical in layout, this will be difficult, since PDFs lack semantic information. Can you instead query a DB for the DOI of the paper (getting the DOI via the filename, or via the title of the paper, assuming the title is easy to grab), and use said DOI DB to get the information in a way that's much easier to process? -- Phillip Gawlowski Though the folk I have met, (Ah, how soon!) they forget When I've moved on to some other place, There may be one or two, When I've played and passed through, Who'll remember my song or my face.