Senior Data Acquisition & Document Intelligence Engineer

via Freelancer ·

Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted1 hour ago
Project Title: Senior Data Acquisition & Document Intelligence Engineer

We are looking for a highly experienced Senior Data Acquisition & Document Intelligence Engineer to build, operate, and continuously improve a large-scale production data acquisition and document processing system.

Project Type: Long-term / Ongoing
Engagement: Full-time
Experience Required: 5+ years
Location: On-site / Hybrid

About the Project

We collect large volumes of publicly available documents and structured records from hundreds of external web sources every day.

The sources are highly inconsistent and frequently change their website structure, formats, URLs, APIs, and document layouts. A significant portion of the data also comes from scanned PDFs, images, and other difficult-to-process documents.

We need someone who can take end-to-end ownership of the entire data acquisition pipeline, including web scraping, crawling, document downloading, OCR, data extraction, structuring, validation, quality assurance, monitoring, and ongoing maintenance.

This is not a one-time scraping project. We are looking for someone with proven experience building and maintaining production-grade extraction pipelines at scale.

Key Responsibilities

1. Web Scraping & Data Acquisition

Build and maintain large-scale crawlers across hundreds of websites and external data sources.
Handle dynamic websites, JavaScript rendering, pagination, sessions, cookies, tokens, forms, and file downloads.
Work with technologies such as Python, Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, etc.
Implement rate limiting, retries, backoff, and request management.
Build incremental crawling and change-detection mechanisms.
Detect website/source changes and quickly fix broken extractors.
Ensure data is collected completely, accurately, and on schedule.

2. OCR & Document Processing

Process large volumes of native PDFs, scanned PDFs, images, and other document formats.
Build and optimize OCR pipelines for poor-quality documents.
Handle skewed/rotated pages, faded text, stamps, seals, watermarks, handwriting, and complex layouts.
Apply image preprocessing such as deskewing, denoising, binarization, cropping, and image enhancement.
Extract structured information from complex tables, including merged cells, multi-page tables, borderless tables, and nested headers.
Process multilingual documents, including Indian regional languages.
Benchmark and select appropriate OCR engines based on document type and accuracy.

3. Data Extraction & Structuring

Convert unstructured web/document content into structured datasets.
Design and maintain data schemas.
Use rules, regex, layout-aware extraction, NLP/NER, and LLM-assisted extraction where appropriate.
Normalize dates, amounts, units, names, addresses, identifiers, and other inconsistent values.
Implement deduplication and entity resolution.
Maintain taxonomies, classification logic, and reference/master data.

4. Data Quality & Validation

Build automated quality checks for completeness, accuracy, freshness, consistency, uniqueness, and validity.
Implement validation and anomaly-detection mechanisms.
Maintain gold-standard datasets and benchmark extraction/OCR accuracy.
Create human-in-the-loop review processes for low-confidence results.
Reconcile extracted data against the original source.
Maintain data lineage and audit trails so records can be traced back to the original source document and extraction run.

5. Production Operations

Build and operate reliable, scheduled production pipelines.
Use Airflow, Prefect, Dagster, or similar orchestration tools.
Implement monitoring, logging, alerting, retries, and failure recovery.
Track uptime, throughput, latency, coverage, cost, and data quality.
Use Docker and CI/CD.
Troubleshoot production issues and prevent recurring failures.
Document the system, extraction logic, and operational procedures.
Must-Have Skills
5+ years of hands-on experience in production data extraction, web scraping, document processing, or data acquisition.
Strong Python programming and production engineering practices.
Strong experience with large-scale web scraping/crawling.
Experience with Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, or equivalent.
Deep practical OCR experience with Tesseract, PaddleOCR, EasyOCR, Google Document AI, AWS Textract, Azure Document Intelligence, or similar.
Strong PDF/document processing experience with PyMuPDF, pdfplumber, pdfminer, Camelot, Tabula, or equivalent.
Strong SQL and relational database knowledge.
PostgreSQL or equivalent database experience.
Experience with Airflow, Prefect, Dagster, or similar workflow orchestration.
Experience with AWS and/or GCP.
Experience with Docker, CI/CD, logging, monitoring, and alerting.
Strong understanding of data quality and automated testing.
Experience maintaining and improving production datasets over time.
Understanding of responsible and lawful web data collection, including robots.txt, terms of service, privacy, and applicable data protection requirements.
Good to Have
OpenCV and image processing
NLP / NER
LLM-based document extraction
Multilingual OCR
Indian regional language OCR
Entity resolution / record linkage
Data lineage and provenance
Great Expectations, Soda, Pandera, or similar
dbt and dbt tests
BigQuery
AWS S3, Lambda, Batch
GCP Cloud Storage
Distributed/batch processing
Experience with government, legal, financial, regulatory, or public-sector documents
Ideal Candidate

We are not looking for someone who has only built basic scraping scripts.

The ideal candidate should have experience operating a production system involving:

Hundreds of Sources → Crawling → Document Acquisition → OCR → Data Extraction → Structuring → Validation → QA → Production Dataset

You should understand common failure points such as website changes, broken selectors, missing documents, OCR errors, duplicate records, incorrect table extraction, silent pipeline failures, source downtime, and data quality degradation—and know how to build systems that detect and recover from these issues.

To Apply

Please send your CV along with a brief description of:

The largest scraping/document extraction pipeline you have personally managed.
Number of sources/websites handled.
Approximate daily/monthly volume of documents or records.
Types of documents processed.
Technologies and frameworks used.
Your measured OCR/extraction accuracy, if available.
How you measured and improved accuracy.
One example where a production pipeline failed silently, how you discovered it, and what you changed to prevent it from happening again.

Please provide specific numbers, technologies, and examples wherever possible. We are particularly interested in candidates with hands-on experience operating large-scale, continuously running production data acquisition systems.
python data processing web scraping ocr scrapy data scraping data extraction data management data engineer ocr automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.