PDF Table Text Extraction Needed

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a collection of PDFs that contain well-formed tables; every page follows the same column layout, so column detection should be straightforward. All I need is the raw text from those tables, delivered as clean, line-break-friendly plain-text files. No Excel, no CSV—just a clear text dump I can later edit and reformat on my side.

Because the column structure never changes, I expect a high level of accuracy without manual post-processing. Feel free to use Python with libraries such as Camelot, Tabula, or even custom regex parsing—whatever reliably keeps each row intact.

Deliverables
• A separate .txt file for each PDF (or a single consolidated .txt if that is easier), preserving the original row order.
• A brief note on the method or script you used so I can reproduce or adapt it later.

Before starting, let me know roughly how long the job will take and what approach you plan to use.
python excel web scraping software architecture data extraction data analysis data management
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.