Extract PDF Text to SQL

via Freelancer ·

Budget / Salary$250–750
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a batch of unstructured PDF documents and want every piece of text inside them pulled out and inserted into an SQL database. I do not need the original layout, fonts, or styling—just clean, plain-text content mapped to sensible fields (e.g., file name, page number, extracted text).

Your task is to create a repeatable workflow or script that:
• Reads each PDF, handles multipage files, and copes gracefully with occasional scan-quality variations.
• Outputs the raw text into an SQL-ready structure and populates a database I can query immediately after the job is done (MySQL or PostgreSQL are both fine—I’ll match whatever you choose).
• Includes clear setup instructions so I can run the process again on new PDFs without extra help.

Accuracy of the captured text and a clean SQL import are the key acceptance criteria. If you plan on using libraries such as pdfminer.six, PyPDF2, Tika, or an OCR layer for image-based pages, please note that in your proposal. A concise demo on a small PDF of mine will be the first milestone, followed by full conversion of the remaining files and delivery of any code or scripts you write.
python sql mysql postgresql database programming sqlite data extraction data management
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.