Automate PDF Table Extraction Task
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a growing archive of multi-page PDF reports and need every table inside them turned into clean, structured data automatically. Manual copy-paste is no longer an option; I want a repeatable process or script that can ingest a batch of PDFs, detect each table, and export the results to one consolidated CSV or Excel file with the original document name and page number preserved for traceability.
Python is my typical stack, so feel free to lean on libraries such as Camelot, Tabula-py, pdfplumber, or any tool you prefer—as long as it runs headless on Windows or Linux and can be scheduled from the command line. Accuracy matters more than raw speed: merged cells, multi-line headers, and occasional rotated pages show up in these files and need to be captured correctly. If a table is malformed, I still need the script to flag it so I can review.
Deliverables
• Source code with clear setup instructions
• A short README explaining dependencies and how to run the batch job
• Sample output generated from three of my test PDFs (I will supply them)
Acceptance Criteria
• At least 98 % of rows match the original PDFs when spot-checked
• No manual intervention required once the script starts
• Output formatted as UTF-8 CSV (one file per PDF) plus an optional master file concatenating them all
If your preferred language or extraction method differs, pitch it—flexibility is welcome so long as the end result is fully automated, accurate table data from PDFs.
Python is my typical stack, so feel free to lean on libraries such as Camelot, Tabula-py, pdfplumber, or any tool you prefer—as long as it runs headless on Windows or Linux and can be scheduled from the command line. Accuracy matters more than raw speed: merged cells, multi-line headers, and occasional rotated pages show up in these files and need to be captured correctly. If a table is malformed, I still need the script to flag it so I can review.
Deliverables
• Source code with clear setup instructions
• A short README explaining dependencies and how to run the batch job
• Sample output generated from three of my test PDFs (I will supply them)
Acceptance Criteria
• At least 98 % of rows match the original PDFs when spot-checked
• No manual intervention required once the script starts
• Output formatted as UTF-8 CSV (one file per PDF) plus an optional master file concatenating them all
If your preferred language or extraction method differs, pitch it—flexibility is welcome so long as the end result is fully automated, accurate table data from PDFs.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.