Build Diabetes Risk Predictor - 06/08/2026 08:49 EDT

via Freelancer ·

Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a collection of patient medical records already cleaned and stored as CSV files, and I’m ready to turn them into an accurate diabetes-risk prediction tool. My preference is to work with a Random Forest model because of its balance between interpretability and performance on tabular health data.

Your task is to take these CSV datasets, craft the full machine-learning pipeline, and return a model that can reliably identify individuals at elevated risk for diabetes. Along the way, please document the preprocessing steps, feature engineering choices, and the metrics you use so I can reproduce and fine-tune the work later if needed.

Deliverables I expect:
• Complete, well-commented Python code (preferably in a Jupyter notebook) that ingests the CSV files, performs preprocessing, trains the Random Forest, and outputs predictions.
• A concise report outlining feature importance, validation scores, and any hyper-parameter tuning performed.
• Instructions for running the model on new patient data.

If you have ideas for improving accuracy—such as trying different class-weight strategies or ensemble tweaks—feel free to include them, but please keep the Random Forest as the core approach. I look forward to seeing how you can turn these records into actionable insights.
python data processing machine learning (ml) data mining statistical analysis data science data analysis predictive analytics
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.