Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

arXiv cs.LG Papers

Summary

This systematic review evaluates methodological reliability in machine learning models for early Chronic Kidney Disease prediction, revealing that data leakage inflates reported accuracy by over 15% and that more than 80% of predictors lack stability across studies.

arXiv:2607.11963v1 Announce Type: new Abstract: The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially misleading. A significant drawback is the lack of organized evaluation regarding methodological concerns. Key issues include data leakage, limited access to temporal patient records and inconsistency in reported clinical indicators. This research offers a systematic literature review of existing CKD prediction studies using interpretable machine learning techniques, where nineteen relevant studies were selected via systematic searches across major academic databases. To assess methodological reliability, this study introduces a structured taxonomy of information leakage and a quantitative leakage scoring framework to systematically evaluate reliability across CKD prediction studies. The analysis reveals a strong relationship between leakage and inflated performance. Here, High leakage-studies report an average accuracy of 95.48%, compared to 80.2% for leakage-free studies, reflecting an increase of approximately 15.28%. Furthermore, a cross-study feature stability analysis shows that only a small subset of predictors is consistently reproducible, with over 80% lacking reliability. Overall, the findings suggest that many reported performance improvements stem from methodological limitations rather than true predictive capability.
Original Article
View Cached Full Text

Cached at: 07/15/26, 04:17 AM

# 1. Introduction
Source: [https://arxiv.org/html/2607.11963](https://arxiv.org/html/2607.11963)
Evaluating Reliability in Machine Learning Models for Early

Chronic Kidney Disease Prediction: A Systematic Review of

Data Leakage and Predictor Stability

Mashrul Hossain, Nafesa Kibria, Fahim Shahriar

mashrul16hossain@gmail\.com, nafesa220128@gmail\.com, fahimshahriar1306@gmail\.com

East West University, Dhaka, Bangladesh

ABSTRACT \-The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare\-related computer science\. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially misleading\. A significant drawback is the lack of organized evaluation regarding methodological concerns\. Key issues include data leakage, limited access to temporal patient records and inconsistency in reported clinical indicators\. This research offers a systematic literature review of existing CKD prediction studies using interpretable machine learning techniques, where nineteen relevant studies were selected via systematic searches across major academic databases\. To assess methodological reliability, this study introduces a structured taxonomy of information leakage and a quantitative leakage scoring framework to systematically evaluate reliability across CKD prediction studies\. The analysis reveals a strong relationship between leakage and inflated performance\. Here, High leakage\-studies report an average accuracy of 95\.48%, compared to 80\.2% for leakage\-free studies, reflecting an increase of approximately 15\.28%\. Furthermore, a cross\-study feature stability analysis shows that only a small subset of predictors is consistently reproducible, with over 80% lacking reliability\. Overall, the findings suggest that many reported performance improvements stem from methodological limitations rather than true predictive capability\.

Keywords: Chronic Kidney Disease, Machine Learning, Data Leakage, Metrics, Feature Analysis, Biomarkers

Chronic Kidney Disease \(CKD\) is a progressive condition characterized by gradual decline in kidney function, affecting more than 850 million individuals worldwide\[[2](https://arxiv.org/html/2607.11963#bib.bib1)\]\[[24](https://arxiv.org/html/2607.11963#bib.bib2)\]\. Early detection of CKD is crucial because symptoms often appear only in advanced stages when treatment options quite limited\. Early identification of CKD is therefore essential to delay disease progression, manage associated comorbidities and reduce the long\-term economic burden on healthcare systems\[[8](https://arxiv.org/html/2607.11963#bib.bib5)\]\[[22](https://arxiv.org/html/2607.11963#bib.bib6)\]\. In recent years, machine learning \(ML\) has emerged as a promising approach for analyzing complex clinical datasets and identifying patterns associated with the early onset of CKD that may not be detected through conventional statistical techniques\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\[[20](https://arxiv.org/html/2607.11963#bib.bib3)\]\.

Machine Learning models can evaluate various types of patient data, such as lab results, demographic information and medical history, to assess a person’s likelihood of developing CKD\. These forecasting systems assist healthcare providers in identifying high\-risk patients sooner and facilitating preventive treatment strategies before significant kidney damage occurs\[[8](https://arxiv.org/html/2607.11963#bib.bib5)\]\[[20](https://arxiv.org/html/2607.11963#bib.bib3)\]\. However, numerous critical challenges continue to restrict the real world implementation of ML driven diagnostic systems in healthcare\. A well known problem is the restricted interpretability associated with numerous high\-performing models, such as deep neural networks and complex ensemble techniques\[[8](https://arxiv.org/html/2607.11963#bib.bib5)\]\[[10](https://arxiv.org/html/2607.11963#bib.bib8)\]\. Clinicians frequently require clear justifications for algorithmic predictions to trust them and integrate them into their medical decision making\. Also, methodological limitations such as dataset bias, limited external validation and small sample sizes from individual centers continue to affect the reliability and generalizability of many existing studies\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\[[20](https://arxiv.org/html/2607.11963#bib.bib3)\]\[[32](https://arxiv.org/html/2607.11963#bib.bib11)\]\.

Recent studies on early CKD prediction indicate a gradual shift from traditional statistical approaches to more advanced machine learning methods\. Logistic regression is frequently utilized as a straightforward baseline model, while ensemble methods based on trees such as Random Forest, XGBoost and CatBoost which are typically applied to improve prediction accuracy\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\[[4](https://arxiv.org/html/2607.11963#bib.bib12)\]\[[19](https://arxiv.org/html/2607.11963#bib.bib4)\]\. These models usually depend on several widely cited clinical indicators, such as serum creatinine concentrations, estimated glomerular filtration rate \(eGFR\), hemoglobin concentrations, markers of albuminuria and demographic variables including age or hypertension history\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\[[2](https://arxiv.org/html/2607.11963#bib.bib1)\]\[[33](https://arxiv.org/html/2607.11963#bib.bib13)\]\[[19](https://arxiv.org/html/2607.11963#bib.bib4)\]\. Multiple studies demonstrate exceptionally high diagnostic precision and robust area\-under the curve \(AUC\) metrics while evaluating their models on selected datasets\. Consequently, models that achieve almost flawless outcomes in controlled studies may not function as effectively in real healthcare settings, where data quality, patient diversity and clinical variability are considerably higher\[[20](https://arxiv.org/html/2607.11963#bib.bib3)\]\[[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\.

Despite the rapid advancements in machine learning research for early CKD prediction, several methodological issues remain insufficiently addressed in the literature\. A key issue in CKD detection studies is the information leakage, which occurs when predictors contain information directly tied to the outcome label used during model training\. For example, variables such as eGFR or serum creatinine are sometimes used as predictors even when CKD labels are derived from the same measurements, which can lead to overly optimistic model performance\[[17](https://arxiv.org/html/2607.11963#bib.bib9)\]\. Again, the predictors which is identified as important often differ across studies, suggesting that variability in feature significance is based on feature selection technique and dataset traits\[[17](https://arxiv.org/html/2607.11963#bib.bib9)\]\. Although previous studies on CKD prediction mainly compared model performance, they lack a unified framework to systematically detect information leakage or evaluate cross\-study feature stability\. This limitation highlights the need for a comprehensive evaluation of methodological validity, consistency of predictors and clinical relevance in existing CKD studies\.

Research Questions : • How much does the data leakage influence ML models effectiveness?

• Which clinical features show consistent stability in early CKD prediction studies and which predictors are dataset specific?

Objectives :This study addresses the mentioned challenges by providing a structured analysis of methodological approaches in machine learning based CKD prediction research\. The primary objectives of this study are outlined as follows: • A categorization of information leakage in CKD machine learning study, classifying direct leakage, proxy leakage and temporal leakage to provide more consistent experimental design in future studies\. • Cross\-study feature stability analysis in three steps to identify clinical predictors that remain consistent across multiple populations and those that seems to be dataset\-specific predictors\.

## 2\. Related Works

Detecting of Chronic Kidney Disease early is vital as it can slow down the disease progression\. CKD doesn’t show symptoms in initial stages, so most cases are found in primary care settings\[[8](https://arxiv.org/html/2607.11963#bib.bib5),[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\. Machine learning interpretable frameworks has emerged as a significant tool to sort complex clinical unorganized data and uncover patterns for risk detection that experienced clinicians may miss\[[20](https://arxiv.org/html/2607.11963#bib.bib3),[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\. Early CKD prediction studies relied on classical algorithms, such as Logistic Regression, Support Vector Machines and Decision Trees\[[30](https://arxiv.org/html/2607.11963#bib.bib22),[29](https://arxiv.org/html/2607.11963#bib.bib24)\]\. These models generalizability was severely limited because of small, single\-center datasets and minimal data quality management\[[9](https://arxiv.org/html/2607.11963#bib.bib19),[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\. Specifically, the high performance scores often reported on the UCI benchmark serve as evidence of overfitting, raising doubts about model reliability in real\-world healthcare settings\. In real\-world with real diversity in patients and clinical chaos, these models would likely stumble\.\[[20](https://arxiv.org/html/2607.11963#bib.bib3),[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\. There is methodological agreement regarding the importance of essential biomarkers such as serum creatinine and estimated glomerular filtration rate \(eGFR\), often identified through feature selection methods like Recursive Feature Elimination \(RFE\), Principal Component Analysis \(PCA\) or LASSO\[[6](https://arxiv.org/html/2607.11963#bib.bib48),[1](https://arxiv.org/html/2607.11963#bib.bib14)\]\. Furthermore, common predictors are albuminuria, patient age and hemoglobin levels\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[3](https://arxiv.org/html/2607.11963#bib.bib23)\]\. A core methodological concern is the instability of leading predictors across studies\. These variations often reflect dataset biases and regional demographic differences such as varying environmental factors and health system constraints\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[15](https://arxiv.org/html/2607.11963#bib.bib18)\]\.

A critical limitation identified is the widespread presence of information leakage\. Numerous studies include pathology\-test\-based markers, such as serum creatinine or eGFR, as input features for model training even though those same exact tests are used to define CKD in the first place\[[17](https://arxiv.org/html/2607.11963#bib.bib9)\]\. This practice introduces circular reasoning, where model ends up just matching the abnormal values instead of actually predicting anything useful, which results in misleading performance metrics that render the model redundant for clinical screening\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[29](https://arxiv.org/html/2607.11963#bib.bib24)\]\. To address this issue, recent work focuses on leakage\-safe preprocessing pipelines, ensuring that operations like scaling, outlier capping and class balancing are performed strictly within the training split\[[2](https://arxiv.org/html/2607.11963#bib.bib1)\]\. Also, Performance evaluation has shifted towards metrics that handle imbalanced data better, like AUC\-ROC, F1\-score and Recall\[[5](https://arxiv.org/html/2607.11963#bib.bib7),[17](https://arxiv.org/html/2607.11963#bib.bib9)\]\.

## 3\. Methodology

### 3\.1 Literature Search Strategy

A detailed search was conducted to identify studies on the early detection of Chronic Kidney Disease utilizing machine learning methods\. Searching through many academic databases, such as IEEE Xplore, PubMed and Google Scholar, several papers were selected in the first stage\.

The search utilized different combinations of particular keywords\. Essential terms included: “CKD" or Chronic Kidney Disease” or “CKD early detection” or “machine learning” or “ML”, “clinical indicators” or “clinical attributes”\. These were combined into queries such as \(“CKD” AND “machine learning” AND “early detection”\) and \(“chronic kidney disease” AND “clinical predictors”\)\.

The search focused on research published from 2021 to 2025 to include recent advancements in machine learning\. Additionally, backward and forward citation tracking was conducted to discover relevant studies beyond the original search results\.

### 3\.2 Study Selection Process

The selection of studies was conducted using a systematic method based on the PRISMA framework\. Studies were first collected from various databases employing the specified search strategy\. Duplicate entries were eliminated and the remaining documents were subjected to title and abstract screening to evaluate their relevance\. Potentially relevant studies were then subjected to full\-text review\. During this stage, papers were assessed based on their emphasis on early\-stage CKD detection, application of machine learning methods, inclusion of clinical predictors and incorporation of interpretability or explainability techniques\. A total of 494 records were identified\. After screening, reducing irrelevant studies, 19 studies were selected\. Figure[1](https://arxiv.org/html/2607.11963#S3.F1)is the visual representation of the study selection process\. This structural filtering ensures that only relevant and high\-quality studies are considered\.

![Refer to caption](https://arxiv.org/html/2607.11963v1/x1.png)Figure 1:PRISMA flowchart for study selection process
### 3\.3 Inclusion and Exclusion Criteria

Explicit inclusion and exclusion criteria were defined to maintain consistency during study selection\.

#### 3\.3\.1 Inclusion Criteria

• Focused on ML\-based early CKD prediction • Addressed early stage prediction • Utilized clinical features \(laboratory or demographic data\) • Published in English within \(2021\-2025\)

#### 3\.3\.2 Exclusion Criteria

• Focused only on medical imaging without predictors • Addressed other diseases \(cardiovascular, diabetes, etc\) • Focused on late\-stage or general CKD detection • Lacked clinical predictor analysis

### 3\.4 Datasets Overview

Datasets used in early chronic kidney disease \(CKD\) research range from relatively small hospital\-based clinical records to very large national repositories such as the UK Biobank, which contains health information from more than half

Table 1:Summary of Datasets Used Across Reviewed Studiesa million participants\[[11](https://arxiv.org/html/2607.11963#bib.bib15)\]\. Among the datasets commonly used for benchmarking machine learning models, the UCI Machine Learning Repository remains one of the most widely adopted sources for algorithm evaluation and comparison\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[13](https://arxiv.org/html/2607.11963#bib.bib28),[16](https://arxiv.org/html/2607.11963#bib.bib27)\]\. At the same time, recent studies increasingly emphasize community\-based screening datasets, particularly in South Asian regions, to support early detection strategies in low\-resource settings where access to advanced laboratory testing is limited\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[27](https://arxiv.org/html/2607.11963#bib.bib33)\]\.

Real\-world clinical datasets contain high\-dimensional feature spaces, irregular time intervals between medical observations and substantial class imbalance, where CKD\-positive cases represent a minority of the population\[[5](https://arxiv.org/html/2607.11963#bib.bib7),[13](https://arxiv.org/html/2607.11963#bib.bib28),[25](https://arxiv.org/html/2607.11963#bib.bib26)\]\. Longitudinal Electronic Medical Record \(EMR\) data further highlight the dynamic nature of disease progression\. For instance, evidence suggests that approximately 20\.3% of diabetic patients showing early kidney impairment develop CKD within six months\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\. Most predictive modeling studies rely on datasets containing roughly 24 to 36 clinical attributes, which include demographic factors as well as biochemical markers such as serum creatinine and estimated glomerular filtration rate \(eGFR\)\[[13](https://arxiv.org/html/2607.11963#bib.bib28),[33](https://arxiv.org/html/2607.11963#bib.bib13),[34](https://arxiv.org/html/2607.11963#bib.bib29)\]\.

However, widely used benchmark datasets may not fully represent the demographic diversity of global populations\. As a result, several studies validate their models using independent cohorts collected from countries such as Bangladesh, Korea, China and the United Arab Emirates\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[21](https://arxiv.org/html/2607.11963#bib.bib31),[35](https://arxiv.org/html/2607.11963#bib.bib32),[10](https://arxiv.org/html/2607.11963#bib.bib8)\]\. National health surveys like NHANES are also frequently used because they follow standardized data collection procedures and incorporate additional lifestyle variables, including dietary patterns and sleep habits, which can influence kidney health outcomes\[[33](https://arxiv.org/html/2607.11963#bib.bib13),[34](https://arxiv.org/html/2607.11963#bib.bib29)\]\. Table[1](https://arxiv.org/html/2607.11963#S3.T1)summarizes the datasets used across reviewed studies\.

### 3\.5 Preprocessing Techniques

Data preprocessing is a crucial step in preparing healthcare datasets for machine learning analysis, as raw clinical data often contain inconsistencies and missing values\[[10](https://arxiv.org/html/2607.11963#bib.bib8),[34](https://arxiv.org/html/2607.11963#bib.bib29)\]\. One of the most common challenges involves handling incomplete records, which may arise from random data loss or manual entry errors during clinical documentation\[[16](https://arxiv.org/html/2607.11963#bib.bib27),[13](https://arxiv.org/html/2607.11963#bib.bib28)\]\. To address this issue, researchers frequently apply imputation techniques such as Multivariate Imputation by Chained Equations \(MICE\), K\-nearest neighbor \(KNN\) imputation or median replacement to reduce the influence of missing information while maintaining robustness against outliers\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[13](https://arxiv.org/html/2607.11963#bib.bib28),[33](https://arxiv.org/html/2607.11963#bib.bib13)\]\.

Additionally to handle temporal variability, longitudinal medical records are frequently summarized within a specific observation window using statistical measures such as the mean, quartiles or standard deviation\[[5](https://arxiv.org/html/2607.11963#bib.bib7)\]\. Feature scaling is also routinely performed to ensure consistent data ranges across variables, typically through Min–Max normalization or Z\-score standardization\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[13](https://arxiv.org/html/2607.11963#bib.bib28),[27](https://arxiv.org/html/2607.11963#bib.bib33),[16](https://arxiv.org/html/2607.11963#bib.bib27)\]\. Moreover, techniques for identifying and correcting outliers like interquartile range \(IQR\) are employed to enhance model stability and prediction accuracy\[[27](https://arxiv.org/html/2607.11963#bib.bib33),[13](https://arxiv.org/html/2607.11963#bib.bib28)\]\. Due to the common issue of class imbalance in CKD datasets, methods such as the Synthetic Minority Over\-sampling Technique \(SMOTE\) or random oversampling are frequently employed to equilibrate the distribution between healthy and CKD\-affected cases\[[13](https://arxiv.org/html/2607.11963#bib.bib28),[25](https://arxiv.org/html/2607.11963#bib.bib26),[11](https://arxiv.org/html/2607.11963#bib.bib15)\]\.

### 3\.6 Feature Selection Methods

Feature selection is equally crucial in determining the most informative clinical predictors\. These techniques are typically categorized into statistical filter methods, model\-based wrapper approaches and regularization strategies\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[27](https://arxiv.org/html/2607.11963#bib.bib33)\]\. Tests like Pearson correlation, ANOVA and the Mann–Whitney U test are frequently utilized to identify variables that significantly vary between CKD and non\-CKD populations\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[13](https://arxiv.org/html/2607.11963#bib.bib28),[11](https://arxiv.org/html/2607.11963#bib.bib15)\]\. Wrapper\-based methods like Recursive Feature Elimination with Cross\-Validation \(RFECV\) and Sequential Feature Selection \(SFS\) have demonstrated the ability to discover smaller subsets of features that often outperform models trained on the complete set of variables\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[27](https://arxiv.org/html/2607.11963#bib.bib33)\]\.

Table 2:Leakage Taxonomy Definition

## 4\. Data Leakage Evaluation

The following section provides a structured evaluation of data leakage in early chronic kidney disease \(CKD\) prediction studies\. To ensure consistent evaluation across studies, three types of data leakage are defined using the following criteria : Type 1—A feature is considered Direct Leakage \(DL\) if it is included in diagnostic criteria for defining CKD\. Here, the criteria consist of estimated glomerular filtration rate, serum creatinine and albuminuria\. Including these as predictive features may lead to circular reasoning or may even hamper the model’s performance, because these features directly impact the clinical understanding of CKD\.

Type 2—Proxy Leakage \(PL\) occurs when a feature is not part of the diagnostic criteria but is statistically or clinically redundant with one\. In this study, proxy leakage is identified using a hybrid criterion: \(i\) strong statistical correlation with diagnostic biomarkers \(\|r\|≥0\.85\|r\|\\geq 0\.85\) or \(ii\) clinical or physiological dependency on the same underlying renal function pathway, even when pairwise correlations are moderate due to dataset\-specific factors\. The \(\|r\|≥0\.85\|r\|\\geq 0\.85\) threshold is commonly adopted as a conservative indicator of strong association in clinical datasets\.

Type 3—Temporal leakage occurs when predictors gathered after the outcome event are utilized in training the model\. Examples include medications after diagnosis, dialysis treatments or subsequent interventions\. Since these variables are unavailable at prediction time, their inclusion leads to unrealistic model performance, a phenomenon widely recognized in machine\-learning evaluation pipelines\[[18](https://arxiv.org/html/2607.11963#bib.bib49)\]\. Table[2](https://arxiv.org/html/2607.11963#S3.T2)presents an overview of the definitions and detection criteria for every leakage category utilized in this analysis\.

To quantify the extent of leakage in each study, a unified leakage score is defined as:

Li=w1​D​Li\+w2​P​Li\+w3​T​LiL\_\{i\}=w\_\{1\}DL\_\{i\}\+w\_\{2\}PL\_\{i\}\+w\_\{3\}TL\_\{i\}\(1\)
whereLiL\_\{i\}represents the total leakage score for studyii, whileD​LiDL\_\{i\},P​LiPL\_\{i\}andT​LiTL\_\{i\}denote direct, proxy and temporal leakage components, respectively\.

The individual leakage components are defined as:

D​Li=\{1,if diagnostic features are included0,otherwiseDL\_\{i\}=\\begin\{cases\}1,&\\text\{if diagnostic features are included\}\\\\ 0,&\\text\{otherwise\}\\end\{cases\}\(2\)P​Li=\{0,\|r\|<0\.61,0\.6≤\|r\|<0\.852,\|r\|≥0\.85PL\_\{i\}=\\begin\{cases\}0,&\|r\|<0\.6\\\\ 1,&0\.6\\leq\|r\|<0\.85\\\\ 2,&\|r\|\\geq 0\.85\\end\{cases\}\(3\)A correlation of\|r\|≥0\.85\|r\|\\geq 0\.85corresponds to a Variance Inflation Factor exceeding6\.76\.7\(V​I​F=1/\(1−r2\)VIF=1/\(1\-r^\{2\}\)\), indicating near\-collinear redundancy with diagnostic biomarkers and warranting full proxy leakage assignment \(P​L=2PL=2\)\[[12](https://arxiv.org/html/2607.11963#bib.bib50),[23](https://arxiv.org/html/2607.11963#bib.bib51)\]\. The intermediate range \(0\.60≤\|r\|<0\.850\.60\\leq\|r\|<0\.85,V​I​F=1\.6VIF=1\.6–6\.76\.7\) captures moderate statistical dependence that may not individually constitute leakage but creates meaningful predictive overlap, particularly when multiple such variables are included simultaneously\[[31](https://arxiv.org/html/2607.11963#bib.bib52)\]\(P​L=1PL=1\), consistent with evidence that partial information leakage can still inflate model performance\[[18](https://arxiv.org/html/2607.11963#bib.bib49)\]\. Variables with\|r\|<0\.60\|r\|<0\.60are treated as statistically independent of diagnostic criteria \(P​L=0PL=0\), following conventional interpretations of correlation strength where values nearr=0\.50r=0\.50represent large effects in behavioral and clinical data\[[7](https://arxiv.org/html/2607.11963#bib.bib53)\]\.

T​Li=\{1,when feature timestamp crosses outcome timestamp0,All other caseTL\_\{i\}=\\begin\{cases\}1,&\\text\{when feature timestamp crosses outcome timestamp\}\\\\ 0,&\\text\{All other case\}\\end\{cases\}\(4\)
Based on the relative severity of each leakage type, weights are assigned as:

w1=2,w2=1,w3=2w\_\{1\}=2,\\quad w\_\{2\}=1,\\quad w\_\{3\}=2\(5\)
Thus, the final leakage scoring function is:

Li=2​D​Li\+P​Li\+2​T​LiL\_\{i\}=2DL\_\{i\}\+PL\_\{i\}\+2TL\_\{i\}\(6\)Leakage Classification Rule

To facilitate interpretation, studies are categorized according to their leakage score:

Leakage Category=\{Low Leakage,Li≤2High Leakage,Li≥3\\text\{Leakage Category\}=\\begin\{cases\}\\text\{Low Leakage\},&L\_\{i\}\\leq 2\\\\ \\text\{High Leakage\},&L\_\{i\}\\geq 3\\end\{cases\}\(7\)
![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_2.png)Figure 2:Conventional vs Leakage aware PipelineFigure[2](https://arxiv.org/html/2607.11963#S4.F2)presents a conceptual comparison between conventional and leakage\-aware CKD prediction workflows\. The figure highlights the importance of appropriate feature handling and leakage control to ensure reliable model evaluation\.

Table 3:Performance Comparison under Controlled Leakage InjectionTable 4:Data Leakage Assessment Across Reviewed Studies### 4\.1Effect of Data Leakage on Model Performance

The analysis reveals a systematic relationship between information leakage and excessive model performance in early CKD prediction studies\. By applying the proposed three\-part taxonomy which are Direct Leakage \(DL\), Proxy Leakage \(PL\) and Temporal Leakage \(TL\), a consistent pattern emerges across the reviewed literature\. Studies that incorporate features directly embedded in diagnostic criteria, such as eGFR, serum creatinine or albuminuria, almost always achieve disproportionately high reported accuracies\. This is not surprising, as these variables are not merely predictive signals but are, in fact, definitional components of the outcome itself\. Consequently, models in these settings are not learning latent disease patterns but are effectively reconstructing the diagnostic rule\. This phenomenon is reflected across a large proportion of studies\[[8](https://arxiv.org/html/2607.11963#bib.bib5),[9](https://arxiv.org/html/2607.11963#bib.bib19),[28](https://arxiv.org/html/2607.11963#bib.bib47),[25](https://arxiv.org/html/2607.11963#bib.bib26),[16](https://arxiv.org/html/2607.11963#bib.bib27),[13](https://arxiv.org/html/2607.11963#bib.bib28),[35](https://arxiv.org/html/2607.11963#bib.bib32)\]which receive the maximum leakage score due to the simultaneous presence of direct and proxy leakage\. Importantly, even in the absence of temporal leakage, the coexistence of direct and proxy leakage is sufficient to produce near\-ceiling performance metrics\.

Table[3](https://arxiv.org/html/2607.11963#S4.T3)demonstrates a continuously increasing trend in reported accuracy as the severity of data leakage increases across models\. In the clean configuration \(no leakage\), models achieved accuracy ranging from 92\.50% \(AdaBoost\) to 97\.50% \(LightGBM\)\. Under the proxy leakage condition, where features correlated with diagnostic criteria, all models exhibited measurable performance gains, with AdaBoost showing the most dramatic improvement of \+5\.00% in accuracy and \+6\.29% in F1\-score\. Under the direct leakage condition, where definitional components of CKD were included as predictors, all models achieved near\-perfect or perfect accuracy \(98\.75–100\.00%\) and AUC \(100\.00%\), regardless of their baseline performance\. Here, Figure[3](https://arxiv.org/html/2607.11963#S4.F3)summarizes the overall workflow of the proposed leakage scoring framework\.

![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_3.png)Figure 3:Leakage Scoring Pipeline
### 4\.2Proxy Correlation Undermining Predictive Learning

A detailed examination of proxy leakage shows how machine learning models can generate overly optimistic performance in chronic kidney disease \(CKD\) prediction tasks\. Proxy leakage occurs when non\-diagnostic variables are highly correlated \(\|r\|≥0\.85\|r\|\\geq 0\.85\) with established diagnostic biomarkers or clinically redundant with diagnostic biomarkers which enables the model to implicitly reconstruct the target label without learning independent predictive patterns\. In clinical datasets, variables such as Blood Urea Nitrogen \(BUN\), Cystatin C and hemoglobin often demonstrate significant correlations with primary indicators like estimated glomerular filtration rate \(eGFR\) and serum creatinine\. Although these variables represent different clinical measurements, their statistical connections allow models to avoid authentic predictive learning by leveraging overlapping data\. This phenomenon is demonstrated in studies such as\[[27](https://arxiv.org/html/2607.11963#bib.bib33)\]and\[[21](https://arxiv.org/html/2607.11963#bib.bib31)\], where moderate leakage correlates with high accuracy, suggesting that even limited proxy leakage can greatly inflate reported performance\.

Importantly, proxy leakage in clinical datasets is not always reflected through high pairwise correlation alone\. In several cases \(UCI\-based studies\), individual correlations with diagnostic biomarkers remain moderate \(\|r\|≈0\.5\|r\|\\approx 0\.5–0\.60\.6\), yet multiple physiologically linked variables \(blood urea nitrogen, hemoglobin, packed cell volume\) are simultaneously included\. These variables collectively encode the same underlying renal dysfunction, forming a multivariate proxy structure\. As a result, PL=2 is assigned in such cases to reflect cumulative redundancy, even in the absence of a single dominant correlation\.

![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_4.png)Figure 4:Data Leakage Severity on Reported AccuracyDue to limited public availability of all datasets used in the reviewed studies \(Table[4](https://arxiv.org/html/2607.11963#S4.T4)\), correlation analysis was conducted on representative datasets \(UCI CKD, Enam Medical, NHANES and Abu Dhabi cohorts\)\. For studies utilizing identical or derived datasets, proxy leakage labels were assigned based on these empirical correlations combined with reported feature sets\.

Apart from performance inflation, proxy leakage affects the learning behavior of machine learning models\. Tree\-based ensemble methods, including XGBoost and Random Forest, are particularly sensitive due to their split\-based optimization process\. During training, these models select features that maximize information gain and highly correlated predictors are often interchangeable in this process\. As a result, proxy variables can repeatedly appear in decision splits, reinforcing the same underlying signal across the ensemble\[[10](https://arxiv.org/html/2607.11963#bib.bib8),[11](https://arxiv.org/html/2607.11963#bib.bib15),[16](https://arxiv.org/html/2607.11963#bib.bib27),[21](https://arxiv.org/html/2607.11963#bib.bib31),[27](https://arxiv.org/html/2607.11963#bib.bib33)\]\. This leads to models that appear to learn clinically meaningful relationships but are actually exploiting statistical redundancy\. Consequently, variables such as hemoglobin and blood urea nitrogen may receive high importance even though they primarily reflect renal impairment already encoded by creatinine or eGFR\[[27](https://arxiv.org/html/2607.11963#bib.bib33),[20](https://arxiv.org/html/2607.11963#bib.bib3),[6](https://arxiv.org/html/2607.11963#bib.bib48)\]\.

Rankings of feature importance under proxy leakage might not represent independent risk factors but rather serve as substitute indicators of diagnostic criteria\. Numerous studies recognize hemoglobin and packed cell volume as important predictors\[[6](https://arxiv.org/html/2607.11963#bib.bib48),[20](https://arxiv.org/html/2607.11963#bib.bib3),[27](https://arxiv.org/html/2607.11963#bib.bib33),[28](https://arxiv.org/html/2607.11963#bib.bib47)\], but these factors are physiologically related to kidney function and highly correlated with creatinine\-based metrics, resulting in erroneous conclusions where models seem to indicate causal relationships but actually reflect disease status\. This likewise diminishes external validity, since models developed on datasets with robust internal correlations frequently perform poorly on external cohorts having different correlation patterns\[[6](https://arxiv.org/html/2607.11963#bib.bib48),[2](https://arxiv.org/html/2607.11963#bib.bib1),[14](https://arxiv.org/html/2607.11963#bib.bib25),[28](https://arxiv.org/html/2607.11963#bib.bib47)\]\. This hybrid formulation ensures that proxy leakage is not underestimated in heterogeneous clinical datasets where statistical correlation alone may fail to capture underlying physiological dependencies\.

Table 5:Performance Comparison- •Welch’s t\-test:Mean Diff\. = 15\.32, SE = 3\.85, t = 3\.98,d​f≈7df\\approx 7, p = 0\.005, 95% CI = \[6\.2, 24\.4\]\. Low leakage if score≤2\\leq 2and High leakage if score≥3\\geq 3\.
- •Observation:Statistically significant difference \(p = 0\.005\), suggesting leakage inflates accuracy by \(6\-24\)%\.
- •Note:It should be noted that the relatively small sample size of the leakage\-low group \(n = 5\) limits statistical power\. Therefore, while the observed difference is statistically significant \(p = 0\.005\), the result should be validated in larger study cohorts\.

### 4\.3Statistical Evidence of Leakage\-Induced Inflation

The aggregate performance comparison in Table[5](https://arxiv.org/html/2607.11963#S4.T5)provides clear evidence of leakage\-induced inflation\. This highlights the importance of rigorous leakage auditing before evaluating model performance or reporting results in early CKD prediction\.

The statistical significance \(p = 0\.005\) of Table[5](https://arxiv.org/html/2607.11963#S4.T5)indicates that the observed difference is very unlikely to be explained solely by sampling variability\. The consistent inflation of model performance due to leakage indicates a methodological bias instead of authentic predictive capability\.

Figure[4](https://arxiv.org/html/2607.11963#S4.F4)represents a gradual rise in reported accuracy as the severity of data leakage increases across multiple studies\. Models classified under the high leakage category consistently report inflated performance, indicating potential overestimation of true predictive ability\. In contrast, leakage\-free studies exhibit relatively lower and more varied accuracy values, indicating more authentic model performance\.

Besides statistical significance, the size of the observed difference indicates a considerable practical effect\. An average inflation exceeding 15\.2% implies that models utilizing predictors prone to leakage could greatly exaggerate their diagnostic ability in practical scenarios\. Importantly, most of the studies where leakage was identified fall into the moderate leakage category, indicating that even small amounts of leakage can influence the reported performance of machine learning models, where inconsistent handling of leakage across studies undermines the reliability of cross\-study performance comparisons and further complicates fair benchmarking of models\.These findings collectively indicate that to have a fair comparison between the published performance benchmarks, we first need to audit and standardize how leakage is handled\.

Table 6:Frequency, Stability and Consistency of Clinical Predictors Across Reviewed CKD Studies \(N=19\)Note: Consistency levels are defined as follows — High: 12–19 studies, Moderate: 6–11 studies, Low: 0–5 studies andNi=frequency of predictor​i​across all papers19N\_\{i\}=\\frac\{\\text\{frequency of predictor \}i\\text\{ across all papers\}\}\{19\}\)

## 5\. Cross\-Study Feature Stability Analysis

The growing use of machine learning models to predict Chronic Kidney Disease \(CKD\) has resulted in a fragmented and often inconsistent understanding of key factors of the disease\. Top predictors of CKD differ across studies, making it difficult to detect characteristics that indicate generalizable risk signals\. This inconsistency raises an essential question: Which predictors are actually reliable across different methods and geographies? which result from dataset\-specific traits, methodological biases?

Table 7:Cross\-Dataset Stability of Clinical PredictorsTo address this problem, a systematic cross\-study feature stability analysis was performed on nineteen CKD prediction studies\. Specifically, 28 standardized attributes \(Table[6](https://arxiv.org/html/2607.11963#S4.T6)\) were classified into groups\. This detailed classification allows for a more refined understanding of feature importance and highlights the interdisciplinary nature of CKD risk modeling\.

Each predictor was evaluated based on its frequency of occurrence across the 19 studies and a normalized stability score was computed\. High\-consistency predictors included patient age, blood pressure, glycemic control, hemoglobin/anemia, serum creatinine and albuminuria\. These features cover multiple clinical areas, reinforcing their apparent robustness across diverse datasets and modeling techniques\.

![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_5.png)Figure 5:Stacked Bar of Clinical DomainsFigure[5](https://arxiv.org/html/2607.11963#S5.F5)presents the distribution of predictors across stability levels \(high, moderate, low\) within each clinical domain\. The renal and clinical areas show the highest concentration of predictors overall and contain the majority of high and moderate stability features, reflecting their frequent inclusion in early CKD prediction models\. In contrast, domains such as lifestyle, socioeconomic, dietary and inflammatory factors are composed almost entirely of low\-stability predictors, indicating restricted usage across various studies\. Numerous domains like anthropometric, metabolic and comorbidity factors include just one or two predictors with low stability representation\. In general, the figure shows a significant disparity in feature utilization, with CKD prediction studies focusing on a narrow range of clinical and renal factors\.

Category\-level aggregation further revealed a disproportionate dominance of renal biomarkers in the literature, with the highest cumulative reporting frequency\. This overrepresentation continues even with the well\- documented risk of data leakage linked with diagnostic variables such as serum creatinine and eGFR\. Their frequent inclusion likely reflects their data availability and diagnostic role rather than true predictive value\.

The cross\-dataset stability analysis of Table[7](https://arxiv.org/html/2607.11963#S5.T7)highlights notable differences between predictors derived from studies using the UCI Chronic Kidney Disease Dataset and those evaluated on independent clinical datasets\. Demographic and clinical variables such as patient age and blood pressure demonstrate the highest cross\-dataset stability, indicating strong generalizability across diverse study populations and modeling approaches\. Across multiple studies, Indicators like glycemic and hemoglobin maintain comparatively high stability, indicating their association with CKD prediction\[[17](https://arxiv.org/html/2607.11963#bib.bib9),[21](https://arxiv.org/html/2607.11963#bib.bib31)\]\. In general, the results suggest that only a limited set of predictors demonstrates strong cross\-population stability, highlighting the necessity for wider dataset variety and standardized feature evaluation for upcoming studies\.

Table 8:Agreement Statistics for Feature SelectionInter\-study agreement in feature selection was evaluated using Fleiss’ kappa \(Table[8](https://arxiv.org/html/2607.11963#S5.T8)\)\. Aκ\\kappavalue of 0\.224 was obtained, indicating fair agreement among the examined studies\. This suggests that although a limited number of predictors \(such as age, blood pressure and glycemic indicators\) appear consistently across studies, the majority of features are inconsistent\. The relatively low k value highlights the disjointed nature of feature selection and supports the observation that many predictors are specific to datasets rather than universally reliable\.

In Figure[6](https://arxiv.org/html/2607.11963#S5.F6), the frequency distribution of all 28 predictors across the 19 reviewed studies is shown, ranked by occurrence\. A steep decline in frequency is observed beyond the sixth predictor, with the majority of features appearing in five or fewer studies\. This distribution visually reinforces the skewed nature of feature selection in CKD prediction research, where a small number of predictors dominate while most are inconsistent\.

Table 9:Comparison of feature stability across all studies \(n=19\) and leakage\-aware studies \(n=5\)The pareto curve of Figure[7](https://arxiv.org/html/2607.11963#S5.F7)demonstrates the combined contribution of predictors to total feature usage across different studies\. A considerable share of overall usage is represented by high\-ranking features as it can be seen from the sharp initial rise in the graph\. Specifically, the initial group of predictors accounts for about 80% of total coverage, illustrating a significant concentration effect\. This pattern highlights the dominance of a limited core feature set in CKD prediction models\.

![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_6.png)Figure 6:Feature Stability With Frequencies![Refer to caption](https://arxiv.org/html/2607.11963v1/Figure_7.png)Figure 7:Pareto Analysis of CKD Predictor Usage Across StudiesTable 10:
Top 10 Core CKD Predictors After Dataset and Leakage AdjustmentThe analysis of Table[9](https://arxiv.org/html/2607.11963#S5.T9)illustrates a distinct ranking of clinically reliable predictors across the five leakage\-safe studies\. Patient age and glycemic control emerge as the most stable and universally utilized features, indicating their robust and consistent predictive relevance\. Evidence from broad inclusion and higher stability, further supports blood pressure and gender as fundamental predictor\. In contrast, variables such as BMI and albuminuria show moderate consistency, suggesting situational relevance across datasets\. Several features, including lifestyle factors and hematological markers, show low stability due to limited inclusion or data availability\. Specifically, traditionally robust clinical predictors like serum creatinine and eGFR are often excluded to avoid data leakage\.

Table[10](https://arxiv.org/html/2607.11963#S5.T10)highlights that patient age, glycemic control and blood pressure are strong indicators, maintaining high performance even after leakage adjustment\. Gender, albuminuria and hemoglobin demonstrate moderate to high stability, enhancing their clinical significance\. Conversely, the decreasing stability of serum creatinine and hematological parameters under leakage\-safe conditions suggests a potential reliance on dataset\-specific signals\.

Only a limited number of predictors display both conceptual independence from diagnostic standards and methodological strength\. Across all three evaluation dimensions, patient age, blood pressure and glycemic control consistently achieve the highest average stability, marking the minimum reliable foundation for CKD risk modeling\. Serum creatinine and albuminuria, despite their stability, are directly associated with CKD diagnosis, raising concerns about diagnostic data\-leakage\. A fragmented research landscape is reflected in the dominance of low\-consistency indicators\. Lifestyle, dietary and socioeconomic factors remain largely unexplored, although they offer potential as leakage\-free early indicators\. Future research should focus on external validation of the stable core predictors and investigate neglected modifiable risk factors\. Moreover, Future work should apply standardized methods for feature selection to enhance the robustness and clinical applicability of machine learning models\.

## 6\. Future Directions

Several clear methodological improvements emerge from this review\. Researchers should implement data leakage auditing protocols throughout model development and follow standardized reporting methods aligned with TRIPOD guidelines\. Additionally, the field would benefit from benchmark datasets that clearly differentiate diagnostic biomarkers from predictive risk factors, enabling more reliable comparisons among models\. Future work could also benefit from focusing on simple temporal features, such as eGFR decline rate or variability across multiple measurements, to more effectively reflect disease progression without relying on complex sequential models\. Furthermore, more reliable detection of data leakage, particularly in resource\-limited environments, can be obtained by examining non\-renal risk factors like lifestyle choices, socioeconomic status\.

## 7\. Limitations

This review is subject to several important constraints\. One limitation is that correlation analysis covers only four representative datasets, which limits the ability to verify findings from studies that used inaccessible data\. Geographical bias presents another concern, as eight of the nineteen reviewed studies use the CKD dataset from an Indian clinical center\. Here, proxy leakage labels for the remaining studies rely on reported feature sets and clinical reasoning\. Also, the review excludes relevant recent works as it is limited to studies published up to 2025\. Moreover, temporal leakage could not be formally validated because of insufficient timestamp data provided in the examined studies\. Additionally, most reviewed studies were evaluated on limited or geographically localized datasets, which may reduce real\-world clinical generalizability\.

## 8\. Conclusion

This systematic review analyzed methods in early CKD prediction using machine learning, with a primary focus on information leakage and cross\-study predictor stability\. One of the findings shows exaggerated predictive effectiveness due to the inclusion of variables that are closely aligned with CKD diagnostic criteria\. This study shows that a considerable amount of reported accuracy likely originates from inherent clinical definitions instead of genuine predictive learning\. This finding is demonstrated by applying a structured taxonomy of direct, proxy and temporal leakage\. Conversely, studies that avoid leakage\-prone features tend to report more moderate but clinically relevant performance, highlighting the necessity of careful feature selection\. Furthermore, the cross\-study analysis indicates a lack of consistency in the importance of low\-leakage predictors, with only a few variables, such as age, blood Pressure and glycemic Control, appearing consistently across multiple datasets\. This variability suggests that model behavior often reflects dataset\-specific characteristics\. Consequently, concerns about generalizability and reproducibility are justified\. Although temporal dynamics were not explicitly analyzed, their potential significance in understanding disease progression is recognized as a key area for future research\. A Pareto curve analysis further supports the findings as it shows that 80% of cumulative feature coverage is achieved with 14 variables, highlighting the diminishing marginal utility of adding extra predictors in low\-leakage CKD models\. Overall, the study highlights the necessity for stricter methodological standards and robust validation frameworks to ensure reliable and clinically significant CKD prediction models\.

Author ContributionsM\.H\. conceived the study, designed the methodology, conducted the systematic review, performed analysis, and interpretation, developed the leakage scoring framework, prepared the figures and tables, and wrote the original manuscript draft\. N\.K\. contributed to the study design, assisted with data interpretation, critically reviewed and revised the manuscript, and provided valuable intellectual input\. F\.S\. contributed to the literature review, assisted with data validation and manuscript revision, and provided constructive feedback throughout the study\. All authors read and approved the final manuscript\.

Corresponding Authormashrul16hossain@gmail\.com \(Mashrul Hossain\)

FundingThere was no funding obtained to conduct this research\.

Competing InterestThe authors declare there were no competing interests\.

Ethics StatementThis study is a systematic review of previously published studies using publicly available and anonymized datasets\. No new human or animal data were collected\. Accordingly, no ethical approval or informed consent was required\. Data AvailabilityNo new datasets were generated or analyzed\. Used dataset is available on UCI Machine Learning Repository\.

Consent for publicationAll authors have read and approved the final manuscript

## References

- \[1\]\(2026\)A systematic review of machine learning methods for chronic kidney disease diagnosis and prediction \(2020–2025\)\.Saudi J Appl Sci Technol2\(1\)\.External Links:[Document](https://dx.doi.org/10.63908/PRV7RM86)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1)\.
- \[2\]M\. R\. Ahmed, M\. A\. Rakib, A\. B\. Shiddik, and M\. S\. Reza\(2025\)Identification of predisposing risk factors for chronic kidney disease and optimizing disease prediction using a stacking machine learning algorithm\.Int J Stat Sci25\(2\),pp\. 1–32\.External Links:[Document](https://dx.doi.org/10.3329/ijss.v25i2.85732)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p2.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.10.9.7.1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.1.1.2.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.17.16.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.19.18.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[3\]M\. S\. Al Huda, E\. Kanon, M\. S\. K\. Pappo, M\. A\. Ali, and N\. Ahmed\(2025\)NefroAI: an explainable and real\-time framework for predicting chronic kidney disease using diverse machine learning models and different feature selection techniques\.IEEE Access14,pp\. 10939–10976\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3649006)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1)\.
- \[4\]I\. A\. Aliet al\.\(2026\)Comparative evaluation of machine learning models for chronic kidney disease diagnosis in resource\-limited healthcare settings using clinically relevant low\-cost biomarkers\.Sci J Publ Health Res Technol,pp\. 35–59\.External Links:[Document](https://dx.doi.org/10.65420/sjphrt.v2i1.64)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1)\.
- \[5\]N\. Aminnejad, M\. Greiver, and H\. Huang\(2025\)Predicting the onset of chronic kidney disease \(ckd\) for diabetic patients with aggregated longitudinal emr data\.PLOS Digit Health4\(1\),pp\. e0000700\.External Links:[Document](https://dx.doi.org/10.1371/journal.pdig.0000700)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p2.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p2.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p3.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.16.15.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.28.27.6.1.1)\.
- \[6\]S\. Boughougal, M\. R\. Laouar, A\. Siam, and S\. Eom\(2025\)An innovative approach for predictive modeling and staging of chronic kidney disease\.Int J Inform Commun Technol\.External Links:[Document](https://dx.doi.org/10.11591/ijict.v14i2.pp684-707)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.12.11.7.1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.20.16.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.16.15.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.17.16.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.19.18.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[7\]J\. Cohen\(2013\)Statistical power analysis for the behavioral sciences\.Routledge\.External Links:[Document](https://dx.doi.org/10.4324/9780203771587)Cited by:[§4\.](https://arxiv.org/html/2607.11963#S4.p6.11)\.
- \[8\]S\. Dutta, R\. Sikder, M\. R\. Islam, A\. Al Mukaddim, M\. A\. Hider, and M\. Nasiruddin\(2024\)Comparing the effectiveness of machine learning algorithms in early chronic kidney disease detection\.J Comput Sci Technol Stud6\(4\),pp\. 77–91\.External Links:[Document](https://dx.doi.org/10.32996/jcsts.2024.6.4.11)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p2.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.17.13.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.17.16.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[9\]J\. Figueroa, P\. Etim, A\. K\. Shibu, D\. Berger, and J\. Levman\(2024\)Diagnosing and characterizing chronic kidney disease with machine learning: the value of clinical patient characteristics as evidenced from an open dataset\.Electronics13\(21\),pp\. 4326\.External Links:[Document](https://dx.doi.org/10.3390/electronics13214326)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.19.15.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.17.16.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.19.18.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[10\]S\. K\. Ghosh and A\. H\. Khandoker\(2024\)Investigation on explainable machine learning models to predict chronic kidney diseases\.Sci Rep14\(1\),pp\. 3687\.External Links:[Document](https://dx.doi.org/10.1038/s41598-024-54375-4)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p2.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.6.5.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.8.4.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1)\.
- \[11\]S\. K\. Ghosh, N\. Widatalla, and A\. H\. Khandoker\(2025\)Machine learning framework for early detection of chronic kidney disease stages using optimized estimated glomerular filtration rate\.IEEE Access\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3565549)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p2.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[§3\.6](https://arxiv.org/html/2607.11963#S3.SS6.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.4.3.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.10.6.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.18.17.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.22.21.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.25.24.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1)\.
- \[12\]J\. F\. Hair, W\. C\. Black, B\. J\. Babin, and R\. E\. Anderson\(2019\)Multivariate data analysis\.Cengage\.Cited by:[§4\.](https://arxiv.org/html/2607.11963#S4.p6.11)\.
- \[13\]M\. E\. Haque, S\. J\. Islam, J\. Maliha, M\. S\. H\. Sumon, R\. Sharmin, and S\. Rokoni\(2025\)Improving chronic kidney disease detection efficiency: fine tuned catboost and nature\-inspired algorithms with explainable ai\.Proc IEEE 14th Int Conf Commun Syst Netw Technol \(CSNT\),pp\. 811–818\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.04262)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p2.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p3.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[§3\.6](https://arxiv.org/html/2607.11963#S3.SS6.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.6.2.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1)\.
- \[14\]J\. He, X\. Wang, P\. Zhu, X\. Wang, Y\. Zhang, J\. Zhao, W\. Sun, K\. Hu, W\. He, and J\. Xie\(2025\)Identification and validation of an explainable early\-stage chronic kidney disease prediction model: a multicenter retrospective study\.EClinicalMedicine84,pp\. 103286\.External Links:[Document](https://dx.doi.org/10.1016/j.eclinm.2025.103286)Cited by:[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.13.12.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.2.2.2.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.20.19.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[15\]H\. Iftikhar, A\. F\. Hashem, M\. Qureshi, and P\. C\. Rodrigues\(2025\)Clinical application of machine learning models for early\-stage chronic kidney disease detection\.Diagnostics15\(20\),pp\. 2610\.External Links:[Document](https://dx.doi.org/10.3390/diagnostics15202610)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1)\.
- \[16\]K\. Jawad, A\. Verma, F\. Amsaad, and L\. Ashraf\(2024\)AI\-driven predictive analytics approach for early prognosis of chronic kidney disease using ensemble learning and explainable ai\.Note:Preprint\. arXiv:2406\.06728\. https://arxiv\.org/abs/2406\.06728 \[Accessed: 1 Jun 2025\]Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p2.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.15.11.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[17\]M\. A\. Kabir, S\. Munira, D\. T\. Azad, S\. M\. Ikram, M\. H\. R\. Sarker, and S\. M\. A\. Hanifi\(2026\)Community\-based early\-stage chronic kidney disease screening using explainable machine learning for low\-resource settings\.Note:Preprint\. arXiv:2601\.01119\. https://arxiv\.org/abs/2601\.01119 \[Accessed: 1 Jun 2025\]Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p4.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p2.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p2.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[§3\.6](https://arxiv.org/html/2607.11963#S3.SS6.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.5.4.7.1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.6.5.7.1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.3.1.1.1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.5.1.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.18.17.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.21.20.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.24.23.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.29.28.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1),[§5\.](https://arxiv.org/html/2607.11963#S5.p6.1)\.
- \[18\]S\. Kapoor and A\. Narayanan\(2023\)Leakage and the reproducibility crisis in machine\-learning\-based science\.Patterns4\(9\),pp\. 100804\.External Links:[Document](https://dx.doi.org/10.1016/j.patter.2023.100804)Cited by:[§4\.](https://arxiv.org/html/2607.11963#S4.p3.1),[§4\.](https://arxiv.org/html/2607.11963#S4.p6.11)\.
- \[19\]S\. M\. Kashani and S\. Z\. B\. Jame\(2023\)Comparing the performance of machine learning models in predicting the risk of chronic kidney disease\.J Arch Mil Med11\(11\),pp\. e145816\.External Links:[Document](https://dx.doi.org/10.5812/jamm-140885)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1)\.
- \[20\]W\. Khalil, K\. Bashir, and M\. Mosadag\(2025\)Early detection of chronic kidney disease \(ckd\) using machine learning algorithms\.East J Comput Sci1\(2\),pp\. 1–9\.External Links:[Document](https://dx.doi.org/10.63496/ejcs.Vol1.Iss2.41)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p2.1),[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.11.10.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.18.14.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[21\]K\. A\. Lee, J\. S\. Kim, Y\. J\. Kim, I\. S\. Goak, H\. Y\. Jin, S\. Park, H\. Kang, and T\. S\. Park\(2025\)A machine learning\-based prediction model for diabetic kidney disease in korean patients with type 2 diabetes mellitus\.J Clin Med14\(6\),pp\. 2065\.External Links:[Document](https://dx.doi.org/10.3390/jcm14062065)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.7.6.7.1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.7.3.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.16.15.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[§5\.](https://arxiv.org/html/2607.11963#S5.p6.1)\.
- \[22\]N\. Nguycharoen\(2024\)Explainable machine learning system for predicting chronic kidney disease in high\-risk cardiovascular patients\.Note:Preprint\. arXiv:2404\.11148\. https://arxiv\.org/abs/2404\.11148 \[Accessed: 1 Jun 2025\]Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1)\.
- \[23\]R\. M\. O’Brien\(2007\)A caution regarding rules of thumb for variance inflation factors\.Qual Quant41\(5\),pp\. 673–690\.External Links:[Document](https://dx.doi.org/10.1007/s11135-006-9018-6)Cited by:[§4\.](https://arxiv.org/html/2607.11963#S4.p6.11)\.
- \[24\]A\. Ortiz, J\. S\. Lees, R\. Torra, V\. S\. Stel, A\. Kramer, and P\. B\. Mark\(2026\)The updated global burden of chronic kidney disease: one death every 20 seconds\.Nephrol Dial Transplant,pp\. gfag040\.External Links:[Document](https://dx.doi.org/10.1093/ndt/gfag040)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p1.1)\.
- \[25\]C\. Paramita and W\. Prasetyaningtyas\(2026\)Enhanced chronic kidney disease prediction using optimized support vector machine with hyperparameter tuning and smote\.Rabit J Teknol Sist Inf Univrab11\(1\),pp\. 964–978\.External Links:[Document](https://dx.doi.org/10.36341/rabit.v11i1.7179)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p3.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.13.9.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.16.15.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1)\.
- \[26\]W\. Pongsittisak and S\. Suraamornkul\(2025\)A simplified machine learning model for predicting reduced kidney function in thai patients with type 2 diabetes: a retrospective study\.J Clin Med14\(13\),pp\. 4735\.External Links:[Document](https://dx.doi.org/10.3390/jcm14134735)Cited by:[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.8.7.7.1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.11.7.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1)\.
- \[27\]C\. N\. E\. Prima and M\. Juhola\(2025\)Early risk factor prediction in chronic kidney disease diagnosis using feature selection and machine learning algorithms\.Methods Inf Med64\(01/02\),pp\. 040–053\.External Links:[Document](https://dx.doi.org/10.1055/a-2797-4380)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p2.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p2.1),[§3\.6](https://arxiv.org/html/2607.11963#S3.SS6.p1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p4.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.16.12.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[28\]N\. G\. Rezk, S\. Alshathri, A\. Sayed, and E\. E\. Hemdan\(2025\)Explainable ai for chronic kidney disease prediction in medical iot: integrating gans and few\-shot learning\.Bioengineering12\(4\),pp\. 356\.External Links:[Document](https://dx.doi.org/10.3390/bioengineering12040356)Cited by:[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.2.1.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2607.11963#S4.SS2.p5.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.21.17.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.11.10.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.13.12.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.14.13.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.15.14.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.17.16.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.19.18.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.5.4.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.9.8.6.1.1)\.
- \[29\]F\. Sanmarchi, C\. Fanconi, D\. Golinelli, D\. Gori, T\. Hernandez\-Boussard, and A\. Capodici\(2023\)Predict, diagnose, and treat chronic kidney disease with machine learning: a systematic literature review\.J Nephrol36,pp\. 1101–1117\.External Links:[Document](https://dx.doi.org/10.1007/s40620-023-01573-4)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1),[§2\.](https://arxiv.org/html/2607.11963#S2.p2.1)\.
- \[30\]K\. V\. Sravya, G\. A\. Reddy, B\. Rajashekar, and T\. N\. Reddy\(2025\)Early prediction of chronic kidney disease\.REST J Data Anal Artif Intell\.External Links:[Document](https://dx.doi.org/10.46632/jdaai/4/1/81)Cited by:[§2\.](https://arxiv.org/html/2607.11963#S2.p1.1)\.
- \[31\]E\. W\. Steyerberg, A\. J\. Vickers, N\. R\. Cook, T\. Gerds, M\. Gonen, N\. Obuchowski, M\. J\. Pencina, and M\. W\. Kattan\(2010\)Assessing the performance of prediction models: a framework for traditional and novel measures\.Epidemiology21\(1\),pp\. 128–138\.External Links:[Document](https://dx.doi.org/10.1097/ede.0b013e3181c30fb2)Cited by:[§4\.](https://arxiv.org/html/2607.11963#S4.p6.11)\.
- \[32\]R\. D\. Suvaris, K\. Nagaiah, P\. Satish, R\. Hussana Johar, E\. Muniyandy, M\. Adusumilli, and K\. Bedair\(2025\)Attention\-enhanced multi\-view graph convolutional network for early prediction of chronic kidney disease\.Int J Adv Comput Sci Appl16\(11\)\.External Links:[Document](https://dx.doi.org/10.14569/IJACSA.2025.0161163)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p2.1)\.
- \[33\]M\. N\. Valencia, J\. Kim, Z\. Abbas, and S\. W\. Lee\(2026\)Early detection of chronic kidney disease in men using lifestyle and demographic indicators: a machine learning approach for primary healthcare settings\.Healthcare14\(3\),pp\. 405\.External Links:[Document](https://dx.doi.org/10.3390/healthcare14030405)Cited by:[§1\.](https://arxiv.org/html/2607.11963#S1.p3.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p3.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.9.5.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.12.11.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.18.17.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.20.19.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.21.20.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.23.22.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.26.25.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.4.3.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.6.5.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1)\.
- \[34\]L\. Zhao, C\. Zhao, Y\. Fu, X\. Wu, X\. Wang, Y\. Wang, and H\. Zheng\(2025\)Oxidative balance score predicts chronic kidney disease risk in overweight adults: a nhanes\-based machine learning study\.Front Nutr12,pp\. 1641496\.External Links:[Document](https://dx.doi.org/10.3389/fnut.2025.1641496)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p3.1),[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[§3\.5](https://arxiv.org/html/2607.11963#S3.SS5.p1.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.3.2.7.1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.14.10.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.10.9.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.18.17.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.20.19.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.21.20.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.23.22.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.24.23.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.25.24.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.26.25.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.3.2.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.8.7.6.1.1)\.
- \[35\]L\. Zou, X\. Wang, Z\. Hou, L\. Sun, and J\. Lu\(2025\)Machine learning algorithms for diabetic kidney disease risk predictive model of chinese patients with type 2 diabetes mellitus\.Ren Fail47\(1\),pp\. 2486558\.External Links:[Document](https://dx.doi.org/10.1080/0886022X.2025.2486558)Cited by:[§3\.4](https://arxiv.org/html/2607.11963#S3.SS4.p4.1),[Table 1](https://arxiv.org/html/2607.11963#S3.T1.6.9.8.7.1.1),[§4\.1](https://arxiv.org/html/2607.11963#S4.SS1.p1.1),[Table 4](https://arxiv.org/html/2607.11963#S4.T4.3.12.8.1.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.16.15.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.2.1.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.22.21.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.27.26.6.1.1),[Table 6](https://arxiv.org/html/2607.11963#S4.T6.5.7.6.6.1.1)\.

Similar Articles