Numerical Data Classification Model
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a structured numerical dataset ready for a clean, reproducible classification workflow. The goal is to build, tune, and evaluate two models—Support Vector Machine and logistic regression—then present the results in a way that lets me decide which approach to take to production.
Here’s the flow I have in mind:
• Pre-process and explore the data (handle missing values, scale where needed, visualise key relationships).
• Implement both classifiers in Python with scikit-learn, using cross-validation and grid/random search for hyper-parameter tuning.
• Produce clear metrics (accuracy, precision-recall, ROC-AUC) and concise plots that compare the two models side-by-side.
• Package everything in a well-commented Jupyter notebook plus a short summary report (PDF or Markdown) that explains findings, chosen parameters, and next steps.
Acceptance criteria
1. Notebook runs end-to-end on my machine with a single cell execution (conda / pip requirements listed).
2. Both Support Vector Machine and logistic regression results are reported using the same validation splits.
3. Code is PEP-8 compliant and functions are logically modular.
4. Summary report highlights why one model might outperform the other and suggests any further improvements.
If you see value in optionally adding a third algorithm such as Random Forest or a Neural Network for comparison, mention it in your proposal—flexibility is welcome as long as the two core models remain the focus.
Preferred stack: Python 3.x, scikit-learn, pandas, NumPy, matplotlib or seaborn.
Here’s the flow I have in mind:
• Pre-process and explore the data (handle missing values, scale where needed, visualise key relationships).
• Implement both classifiers in Python with scikit-learn, using cross-validation and grid/random search for hyper-parameter tuning.
• Produce clear metrics (accuracy, precision-recall, ROC-AUC) and concise plots that compare the two models side-by-side.
• Package everything in a well-commented Jupyter notebook plus a short summary report (PDF or Markdown) that explains findings, chosen parameters, and next steps.
Acceptance criteria
1. Notebook runs end-to-end on my machine with a single cell execution (conda / pip requirements listed).
2. Both Support Vector Machine and logistic regression results are reported using the same validation splits.
3. Code is PEP-8 compliant and functions are logically modular.
4. Summary report highlights why one model might outperform the other and suggests any further improvements.
If you see value in optionally adding a third algorithm such as Random Forest or a Neural Network for comparison, mention it in your proposal—flexibility is welcome as long as the two core models remain the focus.
Preferred stack: Python 3.x, scikit-learn, pandas, NumPy, matplotlib or seaborn.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.