Build Diabetes Risk Predictor - 06/08/2026 08:49 EDT
Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a collection of patient medical records already cleaned and stored as CSV files, and I’m ready to turn them into an accurate diabetes-risk prediction tool. My preference is to work with a Random Forest model because of its balance between interpretability and performance on tabular health data.
Your task is to take these CSV datasets, craft the full machine-learning pipeline, and return a model that can reliably identify individuals at elevated risk for diabetes. Along the way, please document the preprocessing steps, feature engineering choices, and the metrics you use so I can reproduce and fine-tune the work later if needed.
Deliverables I expect:
• Complete, well-commented Python code (preferably in a Jupyter notebook) that ingests the CSV files, performs preprocessing, trains the Random Forest, and outputs predictions.
• A concise report outlining feature importance, validation scores, and any hyper-parameter tuning performed.
• Instructions for running the model on new patient data.
If you have ideas for improving accuracy—such as trying different class-weight strategies or ensemble tweaks—feel free to include them, but please keep the Random Forest as the core approach. I look forward to seeing how you can turn these records into actionable insights.
Your task is to take these CSV datasets, craft the full machine-learning pipeline, and return a model that can reliably identify individuals at elevated risk for diabetes. Along the way, please document the preprocessing steps, feature engineering choices, and the metrics you use so I can reproduce and fine-tune the work later if needed.
Deliverables I expect:
• Complete, well-commented Python code (preferably in a Jupyter notebook) that ingests the CSV files, performs preprocessing, trains the Random Forest, and outputs predictions.
• A concise report outlining feature importance, validation scores, and any hyper-parameter tuning performed.
• Instructions for running the model on new patient data.
If you have ideas for improving accuracy—such as trying different class-weight strategies or ensemble tweaks—feel free to include them, but please keep the Random Forest as the core approach. I look forward to seeing how you can turn these records into actionable insights.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.