From: Kendall Gifford Date: 2013-06-04T05:50:32+09:00 Subject: Re: Question About OCR in Ruby vs. Rails --20cf303dda6285353e04de461d5e Content-Type: text/plain; charset=UTF-8 On Mon, Jun 3, 2013 at 2:28 PM, Kirk Keeter wrote: > Team, > > I'm working on a project that will involve processing 15,000+ complex > financial documents. They are in PDF form. > > Unfortunately, the documents are not available in a non-PDF form -- so I > have to electronically scan the documents and "break them down" into a > database. > Unless the actual data you're trying to extract from the PDF is actually inside an embedded raster image, you don't need OCR but something that can parse the PDF file format. > I'm familiar enough with Rails, that I feel comfortable doing it with the > Rails framework -- but I'm not sure this is a good use of Rails. > So, rails only needs to be involved insofar as you need a web application to wrap or expose this functionality. Otherwise rails is irrelevant to "processing 15,000+ complex financial documents". > Ruby and Javascript are the only programming languages I know, so I'd > either need to somehow do this as a Rails project (with Ruby and > javascript), or as a Ruby project. > > If I do it as a Ruby project (not rails), can you make recommendations > about the best way to go about it? > Searching came up with the pdf-reader gem (https://github.com/yob/pdf-reader) which looks like it'd give you plenty of power to parse and extract the data from your PDF files. Searching also came up with an old (unmaintained?) gem called pdf-toolkit ( https://rubygems.org/gems/pdf-toolkit) that's a wrapper around the pdftk ( http://www.pdflabs.com/tools/pdftk-the-pdf-toolkit/). I'd just play around in an IRB session with these tools, trying to parse out the data from a few representative copies of the documents in question to see what works. Then you could do some trial passes and benchmark them, etc. -- Kendall Gifford zettabyte@gmail.com --20cf303dda6285353e04de461d5e Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable
On Mon, Jun 3, 2013 at 2:28 PM, Kirk Keeter <kirkkeete= r@gmail.com> wrote:

Team,

I'm working on a project that will involve processing 15,000+ comple= x financial documents. =C2=A0They are in PDF form.

Unfortunately, the documents are not available in a non-PDF form -- so I= have to electronically scan the documents and "break them down" = into a database.

Unless the actual data you&= #39;re trying to extract from the PDF is actually inside an embedded raster= image, you don't need OCR but something that can parse the PDF file fo= rmat.
=C2=A0

I'm familiar enough with Rails, that I feel comfortable doing it wit= h the Rails framework -- but I'm not sure this is a good use of Rails.<= br>


So, rails only needs to be in= volved insofar as you need a web application to wrap or expose this functio= nality. Otherwise rails is irrelevant to "processing 15,000+ complex f= inancial documents".
=C2=A0

Ruby and Javascript are the only programming languages I know, so I'= d either need to somehow do this as a Rails project (with Ruby and javascri= pt), or as a Ruby project.

If I do it as a Ruby project (not rails), can you make recommendations a= bout the best way to go about it?<= br>


Searching came = up with the pdf-reader gem (h= ttps://github.com/yob/pdf-reader) which looks like it'd give you pl= enty of power to parse and extract the data from your PDF files. Searching = also came up with an old (unmaintained?) gem called pdf-toolkit (https://rubygems.org/gems/pdf-toolk= it) that's a wrapper around the pdftk (http://www.pdflabs.com/tools/pdftk-the-= pdf-toolkit/). I'd just play around in an IRB session with these to= ols, trying to parse out the data from a few representative copies of the d= ocuments in question to see what works. Then you could do some trial passes= and benchmark them, etc.



--
Kendall Gifford
zettabyte@gmail.com
--20cf303dda6285353e04de461d5e--