PDF to Structured Word Extraction

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted4 hours ago
I have one or more PDFs and need every piece of text pulled out—no images—and placed into a Word document that mirrors the original layout. Headings, paragraphs, lists, tables, footnotes: whatever structure the PDF shows should appear the same way in .docx so I can edit without re-formatting anything myself.

You may use any reliable tool or script (Adobe Acrobat, Python-pdfminer, OCR only if absolutely necessary) as long as the final Word file opens cleanly in the latest Microsoft Word and preserves page order and formatting.

Deliverable
• A single .docx per source PDF containing only the extracted text, laid out to match the PDF’s structure.
• Quality check for missing characters, broken lines, or shifted sections.

I’ll share the PDF(s) once we agree on timing; please let me know how quickly you can turn around a sample page so I can verify the structure is intact.
python data entry pdf latex ocr data extraction adobe acrobat microsoft word
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.