PDF to Structured UTF-8 Text

via Freelancer ·

Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted55 minutes ago
I have a batch of PDF documents that must be turned into clean, well-structured UTF-8 .txt files. Each file contains a mix of narrative text, headings, images, diagrams, tables, and footnotes, so I’m after someone who knows the quirks of PDF extraction inside out.

Here’s exactly what I need you to deliver:
• All text extracted in the correct reading order
• Headings clearly tagged as H1, H2, or H3 so they stand out for future processing
• A placeholder such as [IMAGE: brief description] every time an image, figure, or diagram appears (the PDFs combine text and images)
• Tables converted to plain, readable text rather than skipped or summarised
• Footnotes and endnotes kept as inline text at the point they appear
• No leftover artifacts—page numbers in the middle of paragraphs, broken words, extra line breaks, or stray encoding issues
• Final output: one UTF-8 encoded .txt file for each PDF, mirroring the original structure as closely as possible

Accuracy and attention to detail are critical, especially with heading hierarchy and table formatting. If you’ve handled PDF-to-text conversion before—whether via Acrobat, pdftotext, Python (PyPDF2, pdfminer), OCR tools, or a custom workflow—please tell me a bit about your approach and any relevant projects.

I’ll provide the PDFs once we agree to proceed and would like a short sample from one page first so we can confirm style and structure before you run the full batch.
python technical writing ocr programming scripting data extraction automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.