Screenshot PDF OCR to Database
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a batch of unstructured PDF files in which every page is basically a screenshot that contains text. I need the text pulled out with basic accuracy, organised into a simple, query-friendly data store, and returned so I can run searches and filters on it later.
Here’s how I picture the workflow:
• Run OCR on each screenshot page (Tesseract, Adobe OCR, or any tool you trust) and capture whatever text the engine can reliably identify. Because I only require basic accuracy, occasional missed or garbled characters are acceptable as long as most of the content is captured.
• Push the extracted text into a straightforward structure—CSV, SQLite, or MySQL are all fine. One row per page with fields for file name, page number, and the raw OCR output will meet my needs.
• Deliver the completed database file plus a brief note listing the tool and version you used and any global settings that affect extraction quality (e.g., language packs, dpi adjustments).
The PDFs are ready to share as soon as we start. If anything emerges that makes the screenshots unusually stubborn—blurring, odd fonts, etc.—just flag it and we can decide whether to leave the page blank or accept a rougher pass.
Turnaround within a week would be ideal, but let me know your realistic timeframe when you reply.
Here’s how I picture the workflow:
• Run OCR on each screenshot page (Tesseract, Adobe OCR, or any tool you trust) and capture whatever text the engine can reliably identify. Because I only require basic accuracy, occasional missed or garbled characters are acceptable as long as most of the content is captured.
• Push the extracted text into a straightforward structure—CSV, SQLite, or MySQL are all fine. One row per page with fields for file name, page number, and the raw OCR output will meet my needs.
• Deliver the completed database file plus a brief note listing the tool and version you used and any global settings that affect extraction quality (e.g., language packs, dpi adjustments).
The PDFs are ready to share as soon as we start. If anything emerges that makes the screenshots unusually stubborn—blurring, odd fonts, etc.—just flag it and we can decide whether to leave the page blank or accept a rougher pass.
Turnaround within a week would be ideal, but let me know your realistic timeframe when you reply.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.