TY - GEN
T1 - Federated and Fairness-Aware Learning for Rural Healthcare Risk Prediction Under Data Scarcity
AU - Dhruti, A.
AU - Garg, Saksham
AU - Bhattacharjee, Panchadip
AU - Arukh, Somyajeet
AU - Shet, Nishanth
AU - Gururaj, H. L.
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Centralized machine learning is difficult to implement and susceptible to demographic bias in rural healthcare risk prediction due to fragmented data, small sample sizes, and stringent privacy regulations. In order to facilitate collaborative model training across dispersed rural healthcare sites without exchanging sensitive data, this paper suggests a federated and fairness-aware learning framework. Dirichlet-based partitioning is used to simulate heterogeneous clients using National Family Health Survey data from 707 Indian districts with 14 socioeconomic and healthcare indicators in order to model realistic non-IID conditions. To jointly balance accuracy, equity, and privacy, the suggested federated ensemble integrates gradient-boosted tree models with optimized weighting, empirical differential privacy-inspired noise mechanisms. While maintaining approximately 0.98 of centralized AUC performance (AUC: 0.865 vs. 0.879), the framework reduces socioeconomic disparity gaps and achieves an AUC of 0.865 and an F1-score of 0.705. The introduction of a federated evaluation suite includes uncertainty estimation, feature stability, client contribution attribution, and heterogeneity measurement. Explainability analysis identifies infrastructure access, public healthcare use, and demographic makeup as important risk factors. The findings show a workable approach to fair, comprehensible, and privacy-aware AI risk prediction in simulated federated settings in rural healthcare settings with limited resources.
AB - Centralized machine learning is difficult to implement and susceptible to demographic bias in rural healthcare risk prediction due to fragmented data, small sample sizes, and stringent privacy regulations. In order to facilitate collaborative model training across dispersed rural healthcare sites without exchanging sensitive data, this paper suggests a federated and fairness-aware learning framework. Dirichlet-based partitioning is used to simulate heterogeneous clients using National Family Health Survey data from 707 Indian districts with 14 socioeconomic and healthcare indicators in order to model realistic non-IID conditions. To jointly balance accuracy, equity, and privacy, the suggested federated ensemble integrates gradient-boosted tree models with optimized weighting, empirical differential privacy-inspired noise mechanisms. While maintaining approximately 0.98 of centralized AUC performance (AUC: 0.865 vs. 0.879), the framework reduces socioeconomic disparity gaps and achieves an AUC of 0.865 and an F1-score of 0.705. The introduction of a federated evaluation suite includes uncertainty estimation, feature stability, client contribution attribution, and heterogeneity measurement. Explainability analysis identifies infrastructure access, public healthcare use, and demographic makeup as important risk factors. The findings show a workable approach to fair, comprehensible, and privacy-aware AI risk prediction in simulated federated settings in rural healthcare settings with limited resources.
UR - https://www.scopus.com/pages/publications/105040724457
UR - https://www.scopus.com/pages/publications/105040724457#tab=citedBy
U2 - 10.1109/B-HTC67770.2026.11502232
DO - 10.1109/B-HTC67770.2026.11502232
M3 - Conference contribution
AN - SCOPUS:105040724457
T3 - 2026 IEEE Bangalore Humanitarian Technology Conference, B-HTC 2026
BT - 2026 IEEE Bangalore Humanitarian Technology Conference, B-HTC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 IEEE Bangalore Humanitarian Technology Conference, B-HTC 2026
Y2 - 27 March 2026 through 29 March 2026
ER -