Financial PDF OCR Data Extraction

via Freelancer ·

Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a backlog of invoices, receipts and bank statements, all supplied as searchable and non-searchable PDFs. From each document I only need two categories of information pulled out:

• the dates and amounts that appear on every page
• the full itemised lines (description, quantity, unit price, line total)

Customer names or addresses are not required this time, so the workflow can stay tightly focused on these data points.

Ideally you will set up an OCR pipeline—Tesseract, ABBYY FlexiCapture, Amazon Textract, or a custom Python script with OpenCV—anything you are comfortable with that gets reliable accuracy. The final output should land in a neatly structured CSV or Excel workbook that I can import straight into my accounting software.

Acceptance criteria
• ≥ 98 % field-level accuracy on a random 50-document sample
• Consistent column order: Document ID, Date, Amount, Line Item, Qty, Unit Price, Line Total
• Clear, commented code or repeatable tool configuration so I can rerun the process on new PDFs

Let me know approximately how long you’ll need for an initial batch of 200 documents and which stack you prefer so we can get started right away.
python excel software architecture ocr visual basic for apps opencv data extraction data analysis
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.