Fix Existing n8n Automation OCR & Data Extraction from Image-Based PDFs
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted59 minutes ago
I already have an n8n automation that processes PDF documents received through email. The automation is working, but I am facing an issue with image-based/scanned PDFs.
I receive multiple emails containing PDFs from different document categories. Some PDFs contain selectable text and are processed correctly, but many PDFs are scanned documents or contain the required information as images. My current n8n workflow cannot extract data from these image-based PDFs.
I am looking for an experienced n8n automation developer to update and fix my EXISTING workflow. I do not need the automation rebuilt from scratch.
The updated workflow should:
* Detect whether an incoming PDF is text-based or scanned/image-based.
* Keep the existing extraction process for normal text PDFs.
* Automatically use OCR/image scanning when the PDF does not contain extractable text.
* Extract required information accurately from scanned PDF pages.
* Handle multiple PDF/document categories with different layouts.
* Convert the extracted information into structured data.
* Integrate the OCR process into my existing n8n workflow.
* Handle multi-page PDFs.
* Add proper error handling when OCR or extraction fails.
* Avoid duplicate processing of the same email/document.
* Keep the existing automation functionality working without breaking the current workflow.
The main challenge is not building a new automation. The existing n8n automation is already built and working. I specifically need someone who can diagnose the current PDF extraction limitation and add a reliable OCR/image-based document processing solution to it.
Please apply only if you have experience debugging existing n8n workflows and working with OCR, scanned PDFs, APIs, and document data extraction.
I receive multiple emails containing PDFs from different document categories. Some PDFs contain selectable text and are processed correctly, but many PDFs are scanned documents or contain the required information as images. My current n8n workflow cannot extract data from these image-based PDFs.
I am looking for an experienced n8n automation developer to update and fix my EXISTING workflow. I do not need the automation rebuilt from scratch.
The updated workflow should:
* Detect whether an incoming PDF is text-based or scanned/image-based.
* Keep the existing extraction process for normal text PDFs.
* Automatically use OCR/image scanning when the PDF does not contain extractable text.
* Extract required information accurately from scanned PDF pages.
* Handle multiple PDF/document categories with different layouts.
* Convert the extracted information into structured data.
* Integrate the OCR process into my existing n8n workflow.
* Handle multi-page PDFs.
* Add proper error handling when OCR or extraction fails.
* Avoid duplicate processing of the same email/document.
* Keep the existing automation functionality working without breaking the current workflow.
The main challenge is not building a new automation. The existing n8n automation is already built and working. I specifically need someone who can diagnose the current PDF extraction limitation and add a reliable OCR/image-based document processing solution to it.
Please apply only if you have experience debugging existing n8n workflows and working with OCR, scanned PDFs, APIs, and document data extraction.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.