All projects
Diabetes Risk Modelling and Population Segmentation
Trained and compared XGBoost, Random Forest, and Logistic Regression on a 250,000+ record public health dataset, tuning decision thresholds to maximize recall for high-stakes screening.
Dataset Records250K+
Optimization TargetRecall
SegmentationK-Means & DBSCAN
Pythonscikit-learnXGBoostSHAPpandas
Trained and compared XGBoost, Random Forest, and Logistic Regression on a 250,000+ record public health dataset.
Tuned decision thresholds to maximize recall for high-stakes screening use cases, then applied SHAP for model explainability.
Applied K-Means and DBSCAN for population segmentation and patient profiling.