# Financial Document OCR Specialist Required
Budget / Salary$30–80
TypeFreelance project
LocationRemote
Posted1 hour ago
# Financial Statement Data Extraction Specialist – 800+ Scanned PDFs
I am looking for an experienced **Data Extraction / OCR / Data Engineering specialist** to extract financial data from approximately **800 phone-scanned financial statements** for academic research.
The documents cover **multiple companies and multiple years**, and the final output must be structured as a **panel dataset in Excel** using the template I provide.
### Pilot Test
I have uploaded **5 financial statements as a pilot sample**.
The freelancer should extract the pilot documents into my Excel template exactly as the full project would be processed.
The pilot will be used to evaluate the extraction method and accuracy before awarding the full project.
### What Must Be Extracted
For each company/year:
* All required financial statement numbers
* Correct financial statement line-item/account mapping
* Current-year and comparative-year values
* Company name
* Reporting/financial year
* Auditor / audit firm name
* Audit report date
* **Source PDF filename**
The Excel template I provide will define the required structure and line items.
### Accuracy Requirement
The target is **more than 98% field-level accuracy**.
Each required numeric or metadata field will be checked individually against the original PDF.
For numeric fields, the extracted value must have the correct:
* Number
* Sign
* Year/column
* Account mapping
* Scale/unit
* Company
Errors include wrong values, missing values, wrong signs, wrong years, wrong account mapping, wrong auditor, wrong audit report date, or values assigned to the wrong company.
I will manually verify the pilot to assess the accuracy.
### Quality Control
For traceability, I only require the **source PDF filename** to be linked to each extracted company-year record.
I do **not require** page numbers, bounding-box images, OCR confidence scores, screenshots, or raw OCR text in the final dataset.
The freelancer may use any additional OCR validation or manual review procedures internally to achieve the required accuracy.
Uncertain values should **not be guessed** and should be clearly flagged for review.
### When Applying
Please briefly explain:
* What OCR / Document AI / Python tools you will use
* How you will achieve and measure **98%+ accuracy**
* How you will validate financial numbers and comparative-year columns
* How you will handle low-confidence or unreadable values
* Your experience with similar financial-document extraction
* Your estimated price for approximately **800 PDFs**
**Accuracy, data integrity and consistency are more important than speed.**
I am looking for an experienced **Data Extraction / OCR / Data Engineering specialist** to extract financial data from approximately **800 phone-scanned financial statements** for academic research.
The documents cover **multiple companies and multiple years**, and the final output must be structured as a **panel dataset in Excel** using the template I provide.
### Pilot Test
I have uploaded **5 financial statements as a pilot sample**.
The freelancer should extract the pilot documents into my Excel template exactly as the full project would be processed.
The pilot will be used to evaluate the extraction method and accuracy before awarding the full project.
### What Must Be Extracted
For each company/year:
* All required financial statement numbers
* Correct financial statement line-item/account mapping
* Current-year and comparative-year values
* Company name
* Reporting/financial year
* Auditor / audit firm name
* Audit report date
* **Source PDF filename**
The Excel template I provide will define the required structure and line items.
### Accuracy Requirement
The target is **more than 98% field-level accuracy**.
Each required numeric or metadata field will be checked individually against the original PDF.
For numeric fields, the extracted value must have the correct:
* Number
* Sign
* Year/column
* Account mapping
* Scale/unit
* Company
Errors include wrong values, missing values, wrong signs, wrong years, wrong account mapping, wrong auditor, wrong audit report date, or values assigned to the wrong company.
I will manually verify the pilot to assess the accuracy.
### Quality Control
For traceability, I only require the **source PDF filename** to be linked to each extracted company-year record.
I do **not require** page numbers, bounding-box images, OCR confidence scores, screenshots, or raw OCR text in the final dataset.
The freelancer may use any additional OCR validation or manual review procedures internally to achieve the required accuracy.
Uncertain values should **not be guessed** and should be clearly flagged for review.
### When Applying
Please briefly explain:
* What OCR / Document AI / Python tools you will use
* How you will achieve and measure **98%+ accuracy**
* How you will validate financial numbers and comparative-year columns
* How you will handle low-confidence or unreadable values
* Your experience with similar financial-document extraction
* Your estimated price for approximately **800 PDFs**
**Accuracy, data integrity and consistency are more important than speed.**
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.