CSV EAN Cleanup & Enrichment

via Freelancer ·

Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a CSV database containing approximately 33,000 food products. The database currently has a major data quality issue: the same EAN/barcode is often assigned to multiple completely different products.

For example:

* Peanut Butter - Brand X - EAN 12345
* Rice - Brand X - EAN 12345
* Yogurt - Brand Y - EAN 12345

Obviously, only one of these products can actually correspond to that EAN.

The goal of this project is to verify the database against reliable online sources, identify the correct product for each EAN, remove incorrect records, and deliver a clean database with unique and verified EAN codes.

Scope of Work

1. Verify EAN codes through web scraping / online research

For each EAN code, the contractor should search available online sources to determine which product is actually associated with that barcode.

This may require scraping/searching multiple sources, such as:

* online supermarkets and grocery stores,
* manufacturer websites,
* product databases,
* barcode databases,
* other reliable publicly available sources.

The verification process should determine, whenever possible:

* correct product name,
* correct brand,
* correct EAN/barcode,
* product identity.

2. Resolve duplicate EAN records

Where the same EAN appears on multiple different products, determine which record represents the real product associated with that EAN.

Incorrect records should be removed.

The final database must not contain duplicate EAN codes. One EAN = one product.

3. Remove exact duplicate products

The database may also contain products duplicated independently of the EAN issue.

If the same product name + brand appears multiple times and represents the same product, these duplicate records should also be identified and removed.

The objective is to avoid both:

* duplicate EAN codes,
* duplicate 1:1 product records.

4. Nutrition data enrichment - secondary priority

While verifying products online, I would also like to enrich the correct records with missing nutritional information whenever reliable data is available.

This may include values per 100 g / 100 ml such as:

* calories / energy,
* protein,
* carbohydrates,
* sugars,
* fat,
* saturated fat,
* fiber,
* salt,
* and other nutritional fields already present in the CSV structure.

This is a secondary priority. The main objective is correct EAN identification and database deduplication.

Expected Deliverable

A cleaned CSV file based on the original database where:

* every EAN is assigned to the correct product;
* there are no duplicate EAN codes;
* incorrect products associated with duplicated EANs have been removed;
* exact duplicate products have been removed;
* existing CSV structure/data is preserved where applicable;
* nutritional information is enriched where reliable data can be found.

I am looking for someone experienced in large-scale web scraping, product matching, barcode/EAN data, data cleaning, and deduplication.

Please describe how you would approach the verification process, what data sources you would use, and how you would handle EANs for which no reliable online match can be found.

The database contains approximately 33,000 records, so the solution should be designed for efficient bulk processing rather than manual verification of every individual product.
python data entry excel web scraping web search data scraping beautifulsoup selenium database management
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.