All projects

Diabetes Risk Modelling and Population Segmentation

Trained and compared XGBoost, Random Forest, and Logistic Regression on a 250,000+ record public health dataset, tuning decision thresholds to maximize recall for high-stakes screening.

Dataset Records250K+
Optimization TargetRecall
SegmentationK-Means & DBSCAN
Pythonscikit-learnXGBoostSHAPpandas

Trained and compared XGBoost, Random Forest, and Logistic Regression on a 250,000+ record public health dataset.

Tuned decision thresholds to maximize recall for high-stakes screening use cases, then applied SHAP for model explainability.

Applied K-Means and DBSCAN for population segmentation and patient profiling.