Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease
Summary
The paper presents a two-part study using machine learning on large-scale telehealth data to classify self-reported chronic kidney disease status and identify key psychosocial and medical predictors, achieving balanced accuracy around 72-76% with SHAP analysis for interpretability.
View Cached Full Text
Cached at: 08/19/26, 10:23 AM
# Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease Source: [https://arxiv.org/abs/2608.17174](https://arxiv.org/abs/2608.17174) Authors:[Md\. Atik Shams](https://arxiv.org/search/cs?searchtype=author&query=Shams,+M+A),[David Eisenberg](https://arxiv.org/search/cs?searchtype=author&query=Eisenberg,+D),[Sumaiya Fatema](https://arxiv.org/search/cs?searchtype=author&query=Fatema,+S),[Asma Sultana](https://arxiv.org/search/cs?searchtype=author&query=Sultana,+A),[D\. M Hasibul Islam](https://arxiv.org/search/cs?searchtype=author&query=Islam,+D+M+H),[Junnatul Mawa](https://arxiv.org/search/cs?searchtype=author&query=Mawa,+J),[Anindita Datta](https://arxiv.org/search/cs?searchtype=author&query=Datta,+A),[Nafiya Ahmed](https://arxiv.org/search/cs?searchtype=author&query=Ahmed,+N),[Danastan Tasaouf Mridula](https://arxiv.org/search/cs?searchtype=author&query=Mridula,+D+T),[SK\. Sazid Mahmud](https://arxiv.org/search/cs?searchtype=author&query=Mahmud,+S+S),[Simon Bin Akter](https://arxiv.org/search/cs?searchtype=author&query=Akter,+S+B),[Tanjila Helaly](https://arxiv.org/search/cs?searchtype=author&query=Helaly,+T),[Jorge Fresneda Fernandez](https://arxiv.org/search/cs?searchtype=author&query=Fernandez,+J+F),[Humayera Islam](https://arxiv.org/search/cs?searchtype=author&query=Islam,+H),[Tanmoy Sarkar Pias](https://arxiv.org/search/cs?searchtype=author&query=Pias,+T+S) [View PDF](https://arxiv.org/pdf/2608.17174) > Abstract:Chronic kidney disease \(CKD\) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes\. We present a two\-part study that combines large\-scale telehealth data with advanced machine learning to both classify self\-reported CKD status and identify key drivers of disease\. Using selected features from the Behavioral Risk Factor Surveillance System \(BRFSS 2021: 438,693 samples; BRFSS 2019: 418,268 samples\) and the National Health Interview Survey \(NHIS 2021: 29,482 samples; NHIS 2020: 31,568 samples\), we addressed missing data with nine state\-of\-the\-art imputation methods and mitigated class imbalance via sampling strategies\. Our customized stacked ensemble model achieved balanced accuracy of 72\.56\-76\.12%, with corresponding AUROC scores of 79\.59\-82\.29%\. SHapley Additive exPlanations \(SHAP\) analysis, followed by clinical review, highlighted critical predictors, including regular medical check\-ups, age, blood pressure, and indicators of mental health stress\. These findings deliver a robust and interpretable framework for CKD risk stratification and provide actionable insights into its associated factors\. ## Submission history From: Tanmoy Sarkar Pias \[[view email](https://arxiv.org/show-email/120e3cb1/2608.17174)\] **\[v1\]**Mon, 17 Aug 2026 22:22:13 UTC \(4,074 KB\)
Similar Articles
Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability
This systematic review evaluates methodological reliability in machine learning models for early Chronic Kidney Disease prediction, revealing that data leakage inflates reported accuracy by over 15% and that more than 80% of predictors lack stability across studies.
Calibration, Uncertainty Communication, and Deployment Readiness in CKD Risk Prediction: A Framework Evaluation Study
This study evaluates five machine learning classifiers for chronic kidney disease risk prediction, finding that near-perfect internal performance fails under distribution shift. It emphasizes the need for calibration stability and conformal coverage transfer before clinical deployment.
From Many to Meaningful: Feature-Guided Zero-Shot Chronic Kidney Disease Screening Using Large Language Models
This study proposes a feature-guided zero-shot framework using LLMs for early chronic kidney disease screening, achieving consistent improvements with minimal community-accessible features across heterogeneous datasets.
Transforming Heart Disease Prediction with Advanced Machine Learning Techniques
This research paper compares various machine learning classifiers for heart disease prediction, finding that Support Vector Machine and Simple Cart achieve the best performance on UCI and Kaggle datasets respectively, highlighting ML's potential to aid in early clinical diagnosis.
Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis
This paper develops a multi-dimensional framework to evaluate discrimination, calibration, interpretability, and algorithmic fairness for machine learning-based type 2 diabetes risk prediction models, revealing significant performance degradation under real-world distribution shifts and biases by age and obesity.