Scanned PDF Data Extraction
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted58 minutes ago
I have a collection of forms saved only as scanned‐image PDFs. I need every piece of readable text lifted from those images and placed neatly into an Excel workbook. The end file should let me filter, sort, and analyse the information just as if it had been typed there originally.
Because the source files are images, reliable OCR will be essential. You may use Adobe Acrobat, Tesseract, Python (pandas, openpyxl) or any other toolchain you prefer, as long as the final spreadsheet is clean and ready for immediate use.
Deliverables:
• One .xlsx file containing the extracted text, organised consistently across all forms
• A brief note on the method you used (software or script) so I can reproduce the process if new forms arrive later
Accuracy matters more than speed; I will spot-check the sheet against the original PDFs before sign-off.
Because the source files are images, reliable OCR will be essential. You may use Adobe Acrobat, Tesseract, Python (pandas, openpyxl) or any other toolchain you prefer, as long as the final spreadsheet is clean and ready for immediate use.
Deliverables:
• One .xlsx file containing the extracted text, organised consistently across all forms
• A brief note on the method you used (software or script) so I can reproduce the process if new forms arrive later
Accuracy matters more than speed; I will spot-check the sheet against the original PDFs before sign-off.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.