PDF Heading Extraction Task

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a batch of PDF files that need their text content lifted out and saved as clean .txt files. The most important pieces for me are the headings and titles—those must be captured accurately and appear first in each output file—followed by the remaining text in the order it appears in the source.

Scope
• Open each supplied PDF.
• Copy-paste (or use a reliable extractor) to pull every line of text.
• Make sure each new PDF generates its own plain-text file with the same base filename.
• Headings/titles should be clearly separated—either by an empty line or a simple label—so I can spot them instantly.
• No images, tables, or formatting tags are needed; just raw, readable text.

Acceptance criteria
• One .txt file per PDF, named identically to the source.
• All headings/titles present and readable at the very top.
• No missing, jumbled, or garbled characters.

Speed is appreciated, but accuracy is critical. If you have an automated workflow (Adobe Acrobat, Python pdfminer, OCR for scanned pages, etc.) that meets these requirements, feel free to use it. Let me know how many PDFs you can process per day and when you can start.
php python data entry excel pdf ocr data extraction adobe acrobat
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.