Publish-Ready PDF Text Extraction
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a PDF that I need converted into clean, publish-ready plain text. The end goal is to repurpose the content for online publication, so readability and structure matter just as much as accuracy.
Here’s what I expect:
• Every heading and subheading from the original must be reflected in the .txt file, using a clear hierarchy (e.g., H1, H2, H3).
• Wherever an image or diagram appears, insert an inline placeholder such as “[IMAGE: figure title or brief description]” so I can locate and re-embed the visuals later. No other formatting—no bullet lists, tables, or footnotes—needs to be carried over.
• The text should be stripped of page numbers, headers, footers, and any scanning artifacts, leaving only the main body copy plus the image placeholders mentioned above.
• Final deliverable: a UTF-8 encoded plain-text file that mirrors the PDF’s flow, free of typos and extra spaces.
If you’re comfortable working carefully through PDFs and producing spotless text with a logical heading structure, I’d love your help.
Here’s what I expect:
• Every heading and subheading from the original must be reflected in the .txt file, using a clear hierarchy (e.g., H1, H2, H3).
• Wherever an image or diagram appears, insert an inline placeholder such as “[IMAGE: figure title or brief description]” so I can locate and re-embed the visuals later. No other formatting—no bullet lists, tables, or footnotes—needs to be carried over.
• The text should be stripped of page numbers, headers, footers, and any scanning artifacts, leaving only the main body copy plus the image placeholders mentioned above.
• Final deliverable: a UTF-8 encoded plain-text file that mirrors the PDF’s flow, free of typos and extra spaces.
If you’re comfortable working carefully through PDFs and producing spotless text with a logical heading structure, I’d love your help.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.