Python AI Data Cleaning Automation
Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted3 hours ago
I need a set of well-structured Python scripts that automatically clean and validate raw data so it is ready for downstream analysis. The workflow should use AI-driven techniques where they add value—for example, leveraging an LLM to detect anomalous patterns or flag inconsistent entries more intelligently than simple rule-based checks.
Scope
• Read data from common cloud sources (CSV, JSON or database dump on S3/GCS/Azure Blob).
• Perform thorough cleaning: handle missing values, normalise formats, de-duplicate records, and validate ranges or schema consistency.
• Produce a concise, human-readable report that summarises the cleaning actions taken and highlights rows that still need manual review.
• Package everything into modular, reusable functions (pandas, PySpark or Polars are fine) with clear docstrings and typing.
• Optionally expose the pipeline as a lightweight API or CLI so it can be scheduled in an orchestration tool such as Airflow or executed in a Jupyter notebook.
AI/LLM Angle
If you have experience with LangChain, OpenAI functions or similar, feel free to integrate a step where the model reviews suspicious rows and suggests corrections; however, deterministic rule-based logic must remain the primary validator to keep the pipeline auditable.
Deliverables
1. Clean, commented Python code in a Git-ready repo.
2. A README explaining how to configure cloud credentials and run the pipeline locally or in CI/CD.
3. Sample test cases (pytest) proving the validation rules work.
4. The generated cleaning report in HTML/Markdown for the sample dataset I’ll provide after kickoff.
Acceptance criteria: the test suite passes, the report matches the expected row counts and error rates, and the code deploys without modification on my cloud instance.
I’m aiming for maintainable, production-grade code, so please highlight any best-practice choices you make regarding logging, config management or environment reproducibility.
Scope
• Read data from common cloud sources (CSV, JSON or database dump on S3/GCS/Azure Blob).
• Perform thorough cleaning: handle missing values, normalise formats, de-duplicate records, and validate ranges or schema consistency.
• Produce a concise, human-readable report that summarises the cleaning actions taken and highlights rows that still need manual review.
• Package everything into modular, reusable functions (pandas, PySpark or Polars are fine) with clear docstrings and typing.
• Optionally expose the pipeline as a lightweight API or CLI so it can be scheduled in an orchestration tool such as Airflow or executed in a Jupyter notebook.
AI/LLM Angle
If you have experience with LangChain, OpenAI functions or similar, feel free to integrate a step where the model reviews suspicious rows and suggests corrections; however, deterministic rule-based logic must remain the primary validator to keep the pipeline auditable.
Deliverables
1. Clean, commented Python code in a Git-ready repo.
2. A README explaining how to configure cloud credentials and run the pipeline locally or in CI/CD.
3. Sample test cases (pytest) proving the validation rules work.
4. The generated cleaning report in HTML/Markdown for the sample dataset I’ll provide after kickoff.
Acceptance criteria: the test suite passes, the report matches the expected row counts and error rates, and the code deploys without modification on my cloud instance.
I’m aiming for maintainable, production-grade code, so please highlight any best-practice choices you make regarding logging, config management or environment reproducibility.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.