From: Tamara Temple Date: 2013-06-05T09:40:56+09:00 Subject: Re: Question About OCR in Ruby vs. Rails Kirk Keeter wrote: > I'm working on a project that will involve processing 15,000+ complex financial documents.  They are in PDF form. > > Unfortunately, the documents are not available in a non-PDF form -- so I have to electronically scan the documents and "break > them down" into a database. > > I'm familiar enough with Rails, that I feel comfortable doing it with the Rails framework -- but I'm not sure this is a good use > of Rails. > > Ruby and Javascript are the only programming languages I know, so I'd either need to somehow do this as a Rails project (with > Ruby and javascript), or as a Ruby project. > > If I do it as a Ruby project (not rails), can you make recommendations about the best way to go about it? I don't know if the OCR part is a good use of Rails, etc, but a system to manage all those documents once you've got them OCRed is a spiffy use of Rails. I used to work at a VLC (Very Large Company) in the contracts department, helping put in a new contract management system, and one of the things we had to do was *exactly* this, scanning 10's of thousands of legal documents, some of them exceeding 50-100 pages each. The document management system was completely separate from this, which was a good thing, as we weren't tied to whatever technology for scanning was required. For the scanning bit, we had a handful of really heavy duty Xerox document scanners being driven by dedicated workstations running the full Adobe Acrobat suite. The final scanned, OCRed, documents where stored in a shared drive alongside a file containing their metadata. These were scooped up periodically by a script I wrote that stuffed them into the database of the document management system. There were some serious issues all around, but the biggest headache at the end of the day was getting the people scanning to properly fill out the metadata. Hoo'mahns