Data Evaluation & Benchmark Modeling
Budget / Salary14–30 NZD
TypeFreelance project
LocationRemote
Posted3 hours ago
I need a compact framework that lets me evaluate a data-driven model, benchmark it against reasonable baselines, and then roll those findings into a lightweight “sudo” (pseudo) prediction routine I can run or extend on my own.
The job breaks down into three clear pieces:
1. Design an evaluation pipeline that captures the usual classification/regression metrics and can be adapted to new datasets with minimal code changes.
2. Wire in a benchmarking step so I can see how alternative algorithms or configurations stack up side-by-side—speed and accuracy both matter.
3. Deliver a working prediction script or notebook that reproduces the best-performing setup from the benchmark and outputs predictions in a clean, documented format.
I’m comfortable with either Python (scikit-learn, Pandas, Jupyter) or R (caret, tidyverse) if those are your preferred tools; custom code in another language is acceptable as long as it’s well commented and easy to run.
Acceptance criteria
• Reproducible code base with clear instructions
• Metrics summary and comparison table for each model tested
• Final prediction routine and sample output to confirm correctness
If you’ve built similar evaluation or benchmarking suites before, especially ones that remain readable after the hand-off, I’d love to see an example.
The job breaks down into three clear pieces:
1. Design an evaluation pipeline that captures the usual classification/regression metrics and can be adapted to new datasets with minimal code changes.
2. Wire in a benchmarking step so I can see how alternative algorithms or configurations stack up side-by-side—speed and accuracy both matter.
3. Deliver a working prediction script or notebook that reproduces the best-performing setup from the benchmark and outputs predictions in a clean, documented format.
I’m comfortable with either Python (scikit-learn, Pandas, Jupyter) or R (caret, tidyverse) if those are your preferred tools; custom code in another language is acceptable as long as it’s well commented and easy to run.
Acceptance criteria
• Reproducible code base with clear instructions
• Metrics summary and comparison table for each model tested
• Final prediction routine and sample output to confirm correctness
If you’ve built similar evaluation or benchmarking suites before, especially ones that remain readable after the hand-off, I’d love to see an example.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.