Kidney Res Clin Pract > Volume 45(5); 2026 > Article
Yu, Kim, Joo, Oh, Kim, Lee, Han, Kang, Park, Jeon, and Kim: Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population

Abstract

Background

Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach.

Methods

We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed.

Results

The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934–0.944), significantly exceeding the KFRE (AUROC, 0.879–0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798–0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818–0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification.

Conclusion

The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Graphical abstract

Introduction

Chronic kidney disease (CKD) places a substantial burden on global healthcare systems, rendering the accurate prediction of progression to end-stage renal disease (ESRD) essential for timely clinical intervention. Currently, the Kidney Failure Risk Equation (KFRE), developed by Tangri et al. [1], serves as the clinical gold standard for risk stratification. However, the KFRE relies primarily on cross-sectional data from a single time point, which fails to capture dynamic trajectories of kidney function decline. Furthermore, recent studies using machine learning (ML) suggest that conventional statistical models lack sensitivity and the ability to detect nonlinear risk factors [2,3].
With rapid advances in artificial intelligence (AI), ML and deep learning (DL) models are transforming nephrology. A 2025 review by Li et al. [4] noted that AI models, ranging from traditional decision trees to complex neural networks, handle high-dimensional data that conventional statistics cannot process. Recent evidence supports this potential; Wainstein et al. [5] demonstrated that Random Forest models achieved discrimination comparable or superior to that of the KFRE while incorporating a broader range of clinical variables. Similarly, Moosavi Kashani and Zargar Balaye Jame [6] reported that ensemble learning methods, specifically CatBoost and Random Forest, significantly outperformed artificial neural networks in predicting CKD risk.
Contemporary research is shifting from static predictions toward the incorporation of longitudinal data. Miller and Dwyer [7] emphasized in a recent narrative review that integrating time-dependent variables is crucial for improving model performance. Leung et al. [8] validated this finding, demonstrating that DL algorithms using medical history and prescriptions outperformed the KFRE. Additionally, recent efforts to improve clinical adoption have focused on explainable AI techniques to visualize patient-specific risk factors [2,3]. Barbour et al. [9] demonstrated that clinical models based on KFRE variables (estimated glomerular filtration rate [eGFR], proteinuria, and blood pressure) inadequately predict progression in immunoglobulin A (IgA) nephropathy, requiring disease-specific factors such as MEST (mesangial and endocapillary hypercellularity, segmental sclerosis, and tubular atrophy/interstitial fibrosis) scores and follow-up proteinuria.
In this study, we developed an explainable ML model specific to the Korean population and tailored to tertiary care settings using a large-scale tertiary hospital cohort (Seoul National University Hospital [SNUH]), with external validation in a community-based cohort (Korean Genome and Epidemiology Study [KoGES]) [10]. Our algorithm integrates longitudinal features (e.g., eGFR slope) and uses SHAP analysis to bridge the gap between advanced AI and clinical interpretability.

Methods

Study population and ethical considerations

This study was conducted in accordance with the Declaration of Helsinki. The study protocol was reviewed and determined to be exempt from review by the Institutional Review Board of SNUH (IRB No. E-2507-150-1659, No. E-2512-072-1701), as it used de-identified retrospective data. The requirement for informed consent was waived.

Derivation cohort (Seoul National University Hospital)

The SNUH derivation cohort comprised 28,209 adult patients with CKD, defined according to the 2012 Kidney Disease: Improving Global Outcomes (KDIGO) criteria using the International Classification of Diseases, 10th Revision (ICD-10) code N18.x as a surrogate (eGFR <60 mL/min/1.73 m2 or evidence of kidney damage persisting for >3 months). The data collection period spanned from January 2000 to July 2025. To be included, patients were required to have at least two consecutive eGFR measurements and a minimum of 2 years of continuous observation following their initial diagnosis. Patients were excluded if they had a history of renal replacement therapy (maintenance hemodialysis, peritoneal dialysis, or kidney transplantation) at the time of cohort entry; patients who progressed to ESRD during the observation period were retained and classified as outcome events. Extreme laboratory values were treated as missing and imputed at the variable level rather than triggering patient-level exclusion; accepted ranges and outlier handling logic are detailed in Supplementary Table 1 (available online). The complete patient selection process is illustrated in Fig. 1.

External validation cohort (Korean Genome and Epidemiology Study via Clinical & Omics Data Archive)

The KoGES Ansan-Ansung cohort [10] (n = 10,030, biennial follow-ups) was used for external validation via the Clinical & Omics Data Archive (CODA) secure platform, employing a pooled sliding window across 8 consecutive wave pairs (follow-up 1–2 through 8–9), yielding 3,960 CKD observation-pairs without raw data extraction. The KoGES community cohort did not contain ICD-10 codes; therefore, CKD was phenotyped using KDIGO criteria: baseline eGFR <60 mL/min/1.73 m2 or dipstick proteinuria ≥1+ (as a surrogate for albuminuria), regardless of eGFR, to match the spectrum of the derivation cohort, including early cases manifesting only with proteinuria. For variables unavailable across all KoGES survey waves—specifically serum phosphorus and bicarbonate—SNUH cohort medians were applied uniformly and these variables are reported as ‘Not available’ in Table 1. For variables available only at selected waves (gamma-glutamyl transferase [GGT], total bilirubin, total protein, serum sodium, serum potassium, serum calcium, and serum uric acid), wave-specific measured values were used where available, with SNUH cohort medians applied to waves in which measurements were not collected (footnote ‘b’ in Table 1).

Outcome definition

The primary outcome was a composite of a ≥40% decline in eGFR from baseline or incident ESRD (defined as renal replacement therapy [RRT] initiation or transplantation) within 2 years. ESRD was defined by procedural codes rather than solely by an eGFR <15 mL/min/1.73 m2 to better reflect real-world practice. Participants with a baseline eGFR <15 mL/min/1.73 m2 without prior RRT were included, with outcomes defined as subsequent RRT or a further 40% decline in eGFR; the KoGES validation used identical criteria within the sliding windows. The primary outcome was defined as a ≥40% decline in the eGFR from baseline, while the secondary outcome was a ≥50% decline in eGFR.

Clinical variables and feature engineering

Standardization via Observational Medical Outcomes Partnership Common Data Model

To ensure interoperability and standardization, SNUH electronic health records were converted to the Observational Medical Outcomes Partnership Common Data Model (OMOP-CDM) version 5.3. We used standard OMOP Concept IDs to map local hospital codes to international standard vocabularies: SNOMED-CT for condition diagnoses, RxNorm for medication prescriptions, and LOINC for laboratory measurements (Supplementary Table 2, available online).

Data preprocessing and quality control

Data quality control involved a rigorous inspection of feature distributions. Biologically plausible extremes—such as eGFR >120 mL/min/1.73 m2 (reflecting glomerular hyperfiltration)—were retained to preserve real-world complexity, leveraging the inherent robustness of tree-based ensemble ML models against outliers (Supplementary Fig. 1, available online). Conversely, physiologically impossible values were excluded prior to model training, including potassium >10 mmol/L, hemoglobin <3 g/dL, and eGFR slope values outside the biologically plausible range of ±60 mL/min/1.73 m2 per year (n = 6, 0.02%). Unit discrepancies were corrected where identified (e.g., urine albumin-to-creatinine ratio [UACR] values <0.5 assumed to be in g/g were converted by multiplying by 1,000; values >20,000 mg/g were excluded).
Remaining missing values were addressed by median imputation (scikit-learn SimpleImputer) (Supplementary Table 1, available online), fitted exclusively on the training set and subsequently applied to the internal and external validation sets without refitting, to prevent data leakage. Key missing rates in the derivation cohort were as follows: body mass index (63.7%), UACR (39.8%), urine microalbumin (38.4%), glycated hemoglobin (HbA1c; 18.5%), GGT (32.6%), high-density lipoprotein (9.1%), urine creatinine (8.9%), and triglycerides (8.7%). All baseline statistics presented in Table 1 reflect the finalized dataset following completion of these quality control procedures.

Feature engineering

In the SNUH derivation cohort, serum creatinine was measured using the IDMS-traceable Jaffe method (Hitachi 7600, Roche Diagnostics); IDMS traceability was fully implemented from January 2011 in accordance with National Kidney Disease Education Program recommendations. Pre-2011 measurements may carry a minor systematic offset (approximately 0.1–0.2 mg/dL). The eGFR was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) 2021 equation. In the KoGES external validation cohort, creatinine was measured using the Jaffe method on ADVIA analyzers (Siemens) without IDMS traceability; the CKD-EPI 2009 equation was applied consistent with KoGES analytic guidelines.
• eGFR slope: Calculated as a simple 2-point difference—(eGFR at index date – eGFR at most recent prior visit)/inter-measurement interval (years)—without regression fitting. Critically, only the index-date and prior-visit eGFR were used; the outcome eGFR (subsequent visit) was never incorporated into slope calculation, precluding data leakage. Patients without a prior measurement (Wave 1 entries) had eGFR slope set to missing and imputed with the SNUH training-set median.
• Drug duration: Cumulative exposure to renin-angiotensin-aldosterone system (RAAS) inhibitor (74.8% of derivation cohort) and statins (71.1%) was calculated as days from first prescription to index date, included as continuous features reflecting disease chronicity.

Handling external cohort limitations

In the KoGES external validation cohort, most clinical and laboratory variables aligned well with those of the derivation cohort. A key structural difference from SNUH is the fixed biennial (2-year) measurement interval in KoGES, compared with variable and more frequent laboratory assessments in the tertiary care setting. The sliding-window design across eight consecutive wave pairs (Waves 1–2 through 8–9) was specifically adopted to accommodate this fixed interval, with each wave-pair defining a baseline observation and a standardized 2-year follow-up window; eGFR slope was derived from consecutive wave measurements accordingly. Several variables were unavailable across all KoGES waves and were imputed with SNUH training-set medians: serum bicarbonate (24.0 mEq/L; required for KFRE 8-variable), phosphorus (3.3 mg/dL), sodium (139.7 mmol/L), potassium (4.6 mmol/L), and calcium (9.2 mg/dL; measured at Waves 1 and 4 only). Uric acid, GGT, total bilirubin, total protein, urine creatinine, and urine microalbumin were available in limited waves only and imputed accordingly. For albuminuria, KoGES dipstick results were semi-quantitatively mapped to UACR (negative/trace = 15; 1+ = 30; 2+ = 100; 3+ = 300; 4+ = 1,000 mg/g), consistent with KDIGO staging. Full preprocessing details and accepted ranges are provided in Supplementary Table 1 (available online).

Model development and statistical analysis

We developed a soft-voting ensemble ML model integrating XGBoost, LightGBM, CatBoost, and Random Forest to combine boosting efficiency with bagging diversity. The derivation cohort was randomly divided into training (70%) and internal validation (30%) sets. Stratified random sampling was used so that the primary outcome remained consistent across both subsets. Hyperparameters were optimized via Bayesian optimization with stratified 5-fold cross-validation. The model was compared against the standard 4-variable and 8-variable KFRE using established coefficients [1]. Performance was evaluated using the area under the receiver operating characteristic curve (AUROC, with 95% confidence intervals [CIs]), the area under the precision-recall curve (AUPRC), calibration plots, and decision curve analysis (DCA) to assess clinical utility. All analyses were conducted using Python (version 3.9) with the following key libraries: scikit-learn (version 1.3) for model development and imputation, XGBoost (version 1.7), LightGBM (version 3.3), CatBoost (version 1.2), SHAP (version 0.43) for feature importance analysis, and lifelines (version 0.27) for survival analysis utilities. Pairwise AUROC comparisons between the ensemble ML model and KFRE equations were performed using the DeLong nonparametric test. A prespecified sensitivity analysis was conducted using the secondary composite outcome (≥50% eGFR decline or ESRD), with results reported in Supplementary Table 3 (available online).

Model quality control using calibration plots

Calibration was assessed via visual plots and slope/intercept analysis, excluding the Hosmer-Lemeshow test given its sensitivity to large sample sizes. This study adheres to the TRIPOD+AI reporting guidelines.

Results

Baseline characteristics

The study population comprised 28,209 patients with incident CKD in the SNUH derivation cohort and 3,960 in the KoGES external validation cohort (Fig. 1). The SNUH cohort was randomly divided into a training set (n = 19,746) and an internal validation set (n = 8,463) using stratified random sampling based on the primary outcome. Baseline characteristics were largely balanced between the two subsets (Table 1); minor statistically significant differences observed for serum albumin (4.1 ± 0.5 g/dL vs. 4.0 ± 0.6 g/dL, p = 0.03), total protein (7.0 ± 0.7 g/dL vs. 7.0 ± 0.7 g/dL, p = 0.01), and the glomerulonephritis (GN) subgroup proportion (3.2% vs. 3.7%, p = 0.03) are attributable to the sensitivity of large-sample hypothesis testing to small absolute differences and are not considered clinically meaningful.
Compared with the KoGES cohort, the SNUH cohort had more advanced kidney disease at baseline (mean eGFR, 42.2 ± 27.2 mL/min/1.73 m2 vs. 56.7 ± 13.9 mL/min/1.73 m2), higher proteinuria burden (median UACR, 30 mg/g vs. 15 mg/g), and substantially higher rates of RAAS inhibitor use (74.8% vs. 7.2%) and statin use (71.1% vs. 10.7%), reflecting the tertiary care vs. community-based nature of the respective cohorts. The primary outcome (≥40% eGFR decline or ESRD within 2 years) occurred in 34.3% of SNUH patients and 1.7% of KoGES participants, consistent with the expected difference in disease severity between a hospital-based CKD population and a general community cohort.

Model performance and comparison with the Kidney Failure Risk Equation

In the internal validation cohort (SNUH), the ensemble ML model demonstrated high discriminative performance (AUROC, 0.939; 95% CI, 0.934–0.944), numerically exceeding or matching individual ML models (Table 2, Fig. 2). Notably, the ensemble ML model yielded significantly higher discrimination than the standard KFRE equations (AUROC, 0.884; 95% CI, 0.876–0.893), showing a statistically significant performance gain and a high AUPRC of 0.901. These findings suggest that the ensemble ML model captures complex risk patterns in tertiary care settings more effectively than linear equations. In the external community-based validation (Table 2, Fig. 3), the ensemble ML model maintained robust generalization (AUROC, 0.859; 95% CI, 0.798–0.914), demonstrating performance comparable to that of the KFRE models. While the Random Forest model achieved the highest nominal AUROC of 0.867 with a 95% CI 0.807–0.921 externally, the CIs largely overlapped, supporting the robustness of the ensemble ML approach across diverse clinical environments. DeLong test confirmed that the ensemble ML model significantly outperformed both KFRE equations in the internal validation cohort (vs. KFRE 4-variable: ΔAUROC, +0.061 [95% CI, +0.055 to +0.067], p < 0.001; vs. KFRE 8-variable: ΔAUROC, +0.057 [95% CI, +0.051 to +0.063], p < 0.001) (Supplementary Table 3, available online). In the KoGES external validation cohort, no statistically significant difference was observed between the ensemble model and KFRE (p = 0.35–0.36), consistent with comparable discrimination in the community setting.

Feature importance and interpretability

SHAP analysis confirmed that the model assigned high importance to clinically established risk factors (Fig. 4). Baseline eGFR, serum creatinine, and eGFR slope were identified as the top predictors, underscoring the importance of both initial kidney function and the velocity of decline. Consistent with clinical expectations, lower eGFR, albumin, and hemoglobin, as well as rapid eGFR decline, were associated with increased predicted risk, whereas long-term RAAS inhibitor use showed a protective association. The SHAP dependence plot revealed a distinct nonlinear trajectory in which risk increased substantially below an eGFR of 30 mL/min/1.73 m2 (Fig. 5). Notably, the interaction with UACR was stage-dependent: high UACR significantly amplified risk in patients with preserved eGFR, but this discrimination diminished in advanced CKD (eGFR <30 mL/min/1.73 m2), where low eGFR itself became the dominant prognostic driver regardless of albuminuria.
Of the eight variables comprising the KFRE 8-variable equation, three (eGFR, albumin, and calcium) ranked within the top 10 predictors by SHAP value. Notably, UACR ranked outside the top 20, likely reflecting its high missing rate (39.8%) in the derivation cohort; however, urine creatinine and urine microalbumin—both surrogates of urinary protein excretion—ranked 8th and 9th, respectively, suggesting that the model captured albuminuria-related risk through complementary urinary variables. Serum bicarbonate (TCO2) ranked 16th in overall feature importance, consistent with the established prognostic role of metabolic acidosis in CKD progression and its inclusion in the KFRE 8-variable equation.

Subgroup analysis

To assess the robustness of the ensemble ML model across diverse etiologies, we stratified the SNUH internal validation cohort into four clinically distinct subgroups: diabetic kidney disease (DKD), GN, hypertension, and proteinuria-dominant disease (UACR ≥300 mg/g). The baseline characteristics for each subgroup are detailed in Table 3.
The ensemble ML model demonstrated consistent and high discriminative performance across all subgroups (AUROC >0.91) (Table 4). In contrast to the KFRE models, which exhibited varying degrees of accuracy depending on etiology, the ML model maintained stability.
Notably, the performance gap between the ML and KFRE models was most pronounced in the DKD subgroup. In patients with DKD, the ensemble ML model achieved an AUROC of 0.946 (95% CI, 0.929–0.962), significantly exceeding that of the KFRE 8-variable model (AUROC, 0.866 [95% CI, 0.836–0.893]). This suggests that while linear equations (KFRE) have limited capacity to capture the complex, nonlinear metabolic and vascular risk factors inherent to diabetes, the ML model effectively integrates these heterogeneous signals. In DKD, the ML model captures nonlinear interactions among eGFR slope, HbA1c, albuminuria, and vascular biomarkers—patterns that fixed-coefficient equations cannot represent. In the GN subgroup, both the KFRE (8-variable: AUROC, 0.912 [95% CI, 0.876–0.943]) and ML model (AUROC, 0.941 [95% CI, 0.913–0.966]) performed well. We also considered evaluating rare autoimmune etiologies, including Sjögren syndrome (M35.0x), systemic lupus erythematosus (M32.x), and vasculitis (M30.x) (Supplementary Table 4, available online). However, a dedicated subgroup analysis was not feasible given the limited sample size of these specific conditions within the cohort, which precluded statistically robust validation. A sensitivity analysis using the secondary outcome (≥50% eGFR decline or ESRD) confirmed primary findings; the ensemble ML model maintained discrimination (SNUH AUROC, 0.856; KoGES AUROC, 0.860), while KFRE declined substantially in SNUH (AUROC, 0.667–0.674), supporting ≥40% as the primary endpoint (Supplementary Table 3, available online).

Clinical utility and case study

To illustrate the practical utility of the model, Fig. 6 presents an individual prediction case in which the ML model correctly classified a “rapid progressor” who might have been overlooked by traditional assessments because of preserved baseline function. Despite a relatively stable baseline eGFR (73.2 mL/min/1.73 m2), the model correctly assigned high risk (predicted probability 82%), driven principally by a markedly rapid eGFR slope of –30.6 mL/min/1.73 m2/yr (SHAP, +0.90), hypoalbuminemia (serum albumin, 3.1 g/dL; SHAP, +0.69), anemia (hemoglobin, 9.6 g/dL; SHAP, +0.22), hypocalcemia (calcium, 8.2 mg/dL; SHAP, +0.22), and severe metabolic acidosis (serum bicarbonate, 16 mmol/L; SHAP, +0.14). This demonstrates the capability of the model to capture dynamic risk factors that static scoring systems often miss. Finally, DCA showed that the ensemble ML model provided a higher net benefit than KFRE or “treat-all” strategies across a wide range of threshold probabilities (0.1 to 0.8), supporting its use as a clinical decision support tool in tertiary care settings (Fig. 7).

Quality control using calibration

In the internal validation cohort, the ensemble ML model demonstrated excellent calibration, with the reliability curve closely adhering to the perfect calibration line across the full probability range, indicating accurate risk estimation in the tertiary care setting (Fig. 8A). In the KoGES external validation cohort, at low predicted probabilities (0–0.45), the observed event fraction remained near zero despite increasing predicted risk, reflecting systematic overestimation consistent with the substantially higher event rate of the SNUH derivation cohort (34.3%) relative to the community-based KoGES cohort (1.7%). The substantially lower AUPRC in KoGES (0.489 vs. 0.901 in SNUH) reflects this class imbalance: AUPRC is inherently bounded by outcome prevalence, and a value of 0.489 represents a 28.8-fold improvement over the no-skill baseline (prevalence = 0.017), indicating robust discrimination despite the low event rate. This calibration-in-the-large shift is an expected consequence of applying a hospital-derived model to a general population setting. In the high-probability range (>0.50), the limited number of observations per bin produced instability in the reliability curve (Fig. 8B). These findings suggest that when applying the model in community-based settings, adjustment of the decision threshold or probability rescaling may be warranted to account for the difference in baseline risk between populations.

Discussion

In this study, we developed and externally validated an ensemble ML model specific to the Korean population to predict the rapid progression of CKD. Using the OMOP-CDM, we integrated data from a tertiary hospital (SNUH) and a community-based cohort (KoGES). Our findings revealed a distinct performance pattern dependent on the clinical setting. While the traditional clinical standard, the KFRE, demonstrated high discrimination in the community cohort, our ML model exhibited performance superior to that of the KFRE in the tertiary hospital setting (AUROC, 0.939 vs. 0.879–0.884), while maintaining robust generalizability in the external validation cohort (AUROC, 0.859 [95% CI, 0.798–0.914]). Importantly, the model requires less than 1 msec of inference time per patient, supporting its potential for real-time integration into electronic medical record-based clinical decision support systems. This disparity in performance highlights the structural limitations of linear regression models compared with the flexibility of ML. The KFRE relies on fixed coefficients in which a higher baseline eGFR mathematically dictates a lower predicted risk [1]. Consequently, the equation often fails to identify patients with preserved eGFR as “high risk,” even when rapid decline is imminent. Previous studies have similarly cautioned that the KFRE underestimates risk in complex etiologies such as IgA nephropathy and antineutrophil cytoplasmic antibody-associated vasculitis [8,9]. Our study empirically confirms this limitation in a broader tertiary care population. Our ensemble ML model used nonlinear flexibility [11] to identify these “latent high-risk” patients, particularly in the subgroup with DKD (Table 4). As illustrated in the SHAP dependence plot (Fig. 5), the ML model captured a distinct dispersion pattern at high eGFR levels, using multivariable interactions—specifically eGFR slope and UACR—to distinguish pathologic hyperfiltration from healthy function.
Consistent with the findings of Levey et al. [12] and recent AI reviews [4,7,8], our model identified longitudinal trends, such as the eGFR slope, as dominant predictors. This enables the early detection of patients with rapid functional decline before the onset of kidney failure. Furthermore, the high importance of RAAS inhibitors and albuminuria aligns with established pathophysiology [3,13]. To ensure clinical adoption, we prioritized explainability using SHAP values [14] and observed a higher net benefit in the hospital setting via DCA [15].
A key strength of this study is the implementation of the OMOP-CDM [16], which ensures scalability and facilitates the integration of diverse datasets [1719]. The DCA and calibration plots collectively revealed a setting-dependent pattern of clinical utility. In the SNUH cohort, the ensemble ML model provided clear net benefit over KFRE equations across all clinically plausible thresholds; however, individual ML models exhibited net harm at low threshold probabilities (0–0.3) in the KoGES cohort, consistent with the substantially lower event rate (1.7%) relative to the derivation cohort. The ensemble model’s relative stability near the Treat None line suggests that soft-voting partially mitigates overconfident predictions from individual learners. Calibration analysis similarly demonstrated systematic overestimation at low predicted probabilities in KoGES, reflecting the expected calibration-in-the-large shift when a hospital-derived model is applied to a general population setting; adjustment of the decision threshold or probability rescaling should be considered prior to community-based deployment.
However, this study has some limitations. First, the retrospective design and minimum 2-year follow-up requirement may introduce selection bias, potentially underrepresenting rapid progressors who reached ESRD within 2 years of diagnosis. This may limit generalizability to the broader CKD spectrum. Second, additional drug variables (e.g., loop diuretics, sodium-glucose cotransporter 2 inhibitors) were excluded to avoid overfitting to hospital-specific prescription patterns unavailable in community settings. Third, the low event rate in the KoGES cohort may affect calibration stability [20]. Fourth, the use of real-world data necessitated the inclusion of patients with advanced CKD nearing ESRD or those opting for conservative care, as well as the capture of AKI-on-CKD events. While this introduces heterogeneity, it enhances the sensitivity of the model to rapid, nonlinear declines often missed by traditional models that focus solely on linear progression. Fifth, serum TCO2 (bicarbonate) levels as well as other lab values (phosphorus, urine creatinine, and urine microalbumin) were unavailable in the KoGES cohort and were imputed using the median value from the SNUH derivation cohort. Given that metabolic acidosis is a well-established risk factor associated with declining kidney function, this imputation may have limited the model’s ability to fully leverage the prognostic signal of acid-base status in the external validation. A reduced-feature model using only community-available variables is planned as a future validation step within the SHiNE CDM Consortium expansion. Sixth, serum creatinine measurements prior to IDMS standardization at SNUH (approximately 2000–2010) may carry a minor systematic offset (approximately 0.1–0.2 mg/dL) relative to IDMS-calibrated values. Given the ML model’s reliance on relative longitudinal trajectories rather than absolute creatinine thresholds, this is unlikely to materially affect model performance.
In conclusion, this study presents a robust, explainable, and standardized ensemble ML model for predicting CKD progression in the Korean population. By incorporating longitudinal clinical features and validating across diverse settings—from a tertiary hospital to a community cohort—we demonstrated that this model exceeds the accuracy of current standards derived from Western populations. Prior ML-based CKD progression models [4,7] were predominantly developed in Caucasian cohorts. In Korean and broader East Asian populations, DKD and IgA nephropathy are proportionally more prevalent and metabolic risk profiles differ substantially. Our model, trained in a Korean tertiary care setting and externally validated in a community cohort, directly addresses this gap, particularly demonstrating superior performance in the DKD subgroup. To our knowledge, this is the first study in Korea to externally validate a risk model trained in a hospital setting within a general population cohort. A pivotal strength of this work is its scalability; through the OMOP-CDM framework, the model is designed for rapid expansion across the national healthcare landscape [21] and seamless integration into research consortia such as SHiNE. This infrastructure lays the groundwork for future multi-center “super-cohort” studies. Ultimately, this validated tool paves the way for real-time clinical decision support systems, enabling clinicians to perform precise risk stratification to facilitate early intervention and mitigate the national burden of ESRD.

Notes

Conflicts of interest

A patent based on this work has been filed (application No. 10-2026-0053087, filed March 24, 2026). All authors have no other conflicts of interest to declare.

Funding

This work was supported by the Seoul National University Hospital Research Fund (grant number: 04-2025-2140).

Acknowledgments

This study used biomedical and research resources, including genetic and health information, provided by the Clinical & Omics Data Archive (CODA) and the Korea Disease Control and Prevention Agency, Republic of Korea (approval number: CODA_S2601596-01).

Data sharing statement

The SNUH data are available from the corresponding author upon reasonable request. The KoGES data are available from the Korea Disease Control and Prevention Agency through CODA (https://coda.nih.go.kr) upon approved application.

Authors’ contributions

Conceptualization: HY, YCK

Investigation: HY, YCK, YSK, KWJ, KHO, DKK, HL, SSH, EK, SP

Data curation, Formal analysis, Validation, Visualization, Software: HY

Funding acquisition: HY, BJ

Methodology: HY, YCK

Resources: YCK, YSK, KWJ, KHO, DKK, HL, SSH, EK, SP

Project administration: BJ, YCK

Supervision: YCK

Writing–original draft: HY

Writing–review & editing: All authors

All authors read and approved the final manuscript.

Figure 1.

Study flowchart.

Diagram illustrating the patient selection process for the derivation cohort (SNUH) and external validation cohort (KoGES). The derivation cohort (n = 28,209) was filtered based on age, data availability, and clinical history. The external validation cohort (n = 3,960) was derived using a sliding window approach across 8 data waves.
CKD, chronic kidney disease; eGFR, estimated glomerular filtration rate (mL/min/1.73 m2); Hb, hemoglobin; ICD-10, International Classification of Diseases, 10th Revision; KDIGO, Kidney Disease: Improving Global Outcomes; KoGES, Korean Genome and Epidemiology Study; SNUH, Seoul National University Hospital.
j-krcp-26-055f1.jpg
Figure 2.

Receiver operating characteristic curves in internal validation (Seoul National University Hospital cohort).

Comparison of the ensemble machine learning (ML) model (red) vs. Kidney Failure Risk Equation (KFRE) 4-variable and 8-variable models in the tertiary hospital cohort. The ensemble ML model demonstrated superior discrimination (AUROC, 0.939 [95% CI, 0.934–0.944]) compared to the KFRE models (4-variable: AUROC, 0.879 [95% CI, 0.870–0.887]; 8-variable: AUROC, 0.884 [95% CI, 0.876–0.893]). The diagonal dotted line represents a random classifier (AUROC, 0.50).
AUC, area under the curve; AUROC, area under the receiver operating characteristic curve; CI, confidence interval.
j-krcp-26-055f2.jpg
Figure 3.

Receiver operating characteristic curves in external validation (Korean Genome and Epidemiology Study cohort).

Performance of the ensemble machine learning (ML) model and Kidney Failure Risk Equation (KFRE) models in the community-based external validation cohort (n = 3,960; primary outcome events, 69; event rate, 1.7%). The ensemble ML model maintained robust generalization (AUROC, 0.859 [95% CI, 0.798–0.914]). The KFRE models showed numerically higher AUROC values (4-variable: AUROC, 0.882 [95% CI, 0.818–0.935]; 8-variable: AUROC, 0.889 [95% CI, 0.825–0.941]); however, the DeLong test confirmed no statistically significant difference (p = 0.36 and p = 0.35, respectively), indicating comparable discrimination in the community-based setting. The stepped appearance of the curves reflects the limited number of outcome events (n = 69). The diagonal dotted line represents a random classifier (AUROC = 0.50).
AUC, area under the curve; AUROC, area under the receiver operating characteristic curve; CI, confidence interval.
j-krcp-26-055f3.jpg
Figure 4.

SHAP summary plot.

The top 20 most important features contributing to the model’s predictions, ranked by mean absolute SHAP value. Each dot represents a single patient; red color indicates a high feature value, while blue indicates a low value. Consistent with clinical expectations, baseline estimated glomerular filtration rate (eGFR), serum creatinine, and eGFR slope were identified as the top predictors driving the risk of chronic kidney disease (CKD) progression. Serum bicarbonate (TCO2) appears among the top 20 features, consistent with its role as a surrogate marker for metabolic acidosis severity in CKD.
ACE, angiotensin-converting enzyme; ALT, alanine aminotransferase; ARB, angiotensin-II receptor blocker; AST, aspartate aminotransferase; BMI, body mass index; BUN, blood urea nitrogen; GGT, gamma-glutamyl transferase; Hb, hemoglobin; HbA1c, glycated hemoglobin; SHAP, SHapley Additive exPlanations; TCO2, total carbon dioxide; TG, triglyceride.
j-krcp-26-055f4.jpg
Figure 5.

SHAP dependence plot illustrating the interaction between kidney function and albuminuria.

The plot displays baseline estimated glomerular filtration rate (eGFR, x-axis) against its SHAP value (y-axis), with color representing the urine albumin-to-creatinine ratio (UACR). The nonlinear trajectory highlights a sharp escalation in risk as eGFR falls below 30 mL/min/1.73 m2. In the preserved eGFR range (>60 mL/min/1.73 m2), patients with high UACR (pink/red dots) show increased predicted risk, illustrating the model’s ability to capture pathologic hyperfiltration.
SHAP, SHapley Additive exPlanations.
j-krcp-26-055f5.jpg
Figure 6.

Individual prediction case study (waterfall plot).

A representative “Rapid Progressor” case correctly identified by the ensemble machine learning (ML) model despite preserved baseline kidney function. The waterfall plot illustrates how each feature value shifts the model prediction from the population base value (E[f(X)] = 0.106) to the final high-risk output (f(x) = 1.8; predicted probability = 82%). Red bars indicate risk-increasing contributions; blue bars indicate risk-decreasing (protective) contributions. Despite a relatively preserved baseline estimated glomerular filtration rate (eGFR) of 73.2 mL/min/1.73 m2 (SHAP –0.46, protective), the model correctly assigned high risk driven principally by a markedly rapid eGFR slope of –30.6 mL/min/1.73 m2/yr (SHAP +0.90) and reduced serum albumin of 3.1 g/dL (SHAP +0.69). Low serum bicarbonate (16 mEq/L; SHAP +0.14) further contributed to risk, illustrating how TCO2 captures metabolic acidosis as an independent progression signal.
ACE, angiotensin-converting enzyme; ARB, angiotensin-II receptor blocker; GGT, gamma-glutamyl transferase; SHAP, SHapley Additive exPlanations; TCO2, total carbon dioxide; TG, triglyceride.
j-krcp-26-055f6.jpg
Figure 7.

Decision curve analysis.

Net benefit curves for the ensemble machine learning (ML) model, individual ML models (XGBoost, LightGBM, CatBoost, Random Forest), and Kidney Failure Risk Equation (KFRE) equations (4-variable and 8-variable) across threshold probabilities from 0 to 0.9. The gray dotted line represents the “Treat All” strategy and the black solid line represents the “Treat None” (net benefit = 0) reference. (A) Internal validation (Seoul National University Hospital [SNUH]): The ensemble ML model and individual ML models demonstrate consistently positive net benefit across the full range of clinically relevant threshold probabilities, with all models closely overlapping. Both the Ensemble and component models substantially outperform the KFRE equations, particularly at thresholds above 0.2, where the performance gap widens progressively. The KFRE models fall below the ML models throughout, approaching the Treat None line at higher thresholds. (B) External validation (Korean Genome and Epidemiology Study [KoGES]): In the community-based cohort, the Treat All reference line is substantially attenuated, reflecting the low endpoint prevalence (1.7%). At low threshold probabilities (0–0.3), the individual ML models exhibit net harm—falling below the Treat None line—indicating that their higher sensitivity at this threshold leads to a disproportionate number of false-positive interventions relative to true benefit in a low-prevalence setting. In contrast, the Ensemble model and KFRE equations remain at or marginally above the Treat None line across most thresholds, demonstrating clinical safety. This pattern reflects the low event rate (1.7%) in the KoGES community cohort: the KFRE’s simpler linear structure is adequate for broad population screening, whereas the ML model’s advantage is preserved in the tertiary care setting where outcome prevalence is substantially higher. Beyond threshold probabilities of approximately 0.3, all models converge toward zero net benefit, consistent with the low event rate.
j-krcp-26-055f7.jpg
Figure 8.

Calibration plots.

Assessment of model calibration in the (A) internal Seoul National University Hospital (SNUH) cohort and (B) external Korean Genome and Epidemiology Study (KoGES) cohort. The x-axis represents the mean predicted probability, and the y-axis represents the observed fraction of positives. (A) Internal validation (SNUH): The ensemble machine learning model (red line) tracks closely to the ideal diagonal dotted line, indicating excellent calibration in the hospital setting (Brier score = 0.0894; values closer to 0 indicate better calibration). (B) External validation (KoGES; Brier score = 0.0231): At low predicted probabilities (0–0.45), the observed event fraction remains near zero despite increasing predicted risk, reflecting systematic overestimation consistent with the substantially higher event rate of the SNUH derivation cohort (34.3%) relative to the community-based KoGES cohort (1.7%). This calibration-in-the-large shift is an expected consequence of applying a hospital-derived model to a general population setting. In the high-probability range (>0.50), the limited number of observations per bin produces instability in the reliability curve. These findings suggest that when applying the model in community-based settings, adjustment of the decision threshold or probability rescaling may be warranted to account for the difference in baseline risk between populations.
j-krcp-26-055f8.jpg
j-krcp-26-055f9.jpg
Table 1.
Baseline characteristics of the derivation and external validation cohorts
Characteristic SNUH training set (n = 19,746) SNUH internal validation (n = 8,463) KoGES external validation (n = 3,960) p-value (training vs. internal validation) p-value (SNUH vs. KoGES)
Demographics
 Age (yr) 64.1 ± 15.5 64.2 ± 15.5 66.9 ± 9.4 0.76 <0.001
 Male sex 12,261 (62.1) 5,290 (62.5) 1,672 (40.7) 0.52 <0.001
 Body mass index (kg/m2) 23.9 ± 3.9 24.0 ± 3.9 25.1 ± 3.4 0.22 <0.001
Laboratory measurements
 eGFR (mL/min/1.73 m2) 42.2 ± 27.2 42.1 ± 27.1 56.7 ± 13.9 0.73 <0.001
 Hemoglobin (g/dL) 12.0 ± 2.2 12.0 ± 2.2 13.3 ± 1.6 0.56 <0.001
 HbA1c (%) 6.3 ± 1.2 6.3 ± 1.2 6.1 ± 1.2 0.80 <0.001
 Serum albumin (g/dL) 4.1 ± 0.5 4.0 ± 0.6 4.0 ± 0.2 0.03 <0.001
 Glucose (mg/dL) 118.9 ± 45.7 119.0 ± 45.8 103.9 ± 33.4 0.82 <0.001
 Uric acid (mg/dL) 6.5 ± 2.0 6.5 ± 2.0 5.7 ± 1.7 (n = 1,737)a 0.98 <0.001
 Total cholesterol (mg/dL) 147.3 ± 50.9 147.1 ± 51.6 186.9 ± 40.3 0.79 <0.001
 Triglycerides (mg/dL) 135.8 ± 83.7 135.5 ± 82.0 159.2 ± 108.6 0.74 <0.001
 HDL cholesterol (mg/dL) 47.9 ± 15.1 48.1 ± 15.3 45.4 ± 11.9 0.39 <0.001
 Blood urea nitrogen (mg/dL) 33.0 ± 21.0 33.1 ± 21.2 19.0 ± 6.7 0.60 <0.001
 Serum creatinine (mg/dL) 2.9 ± 2.9 2.9 ± 3.0 1.2 ± 0.5 0.66 <0.001
 Urine creatinine (mg/dL)b 103.0 ± 66.0 102.4 ± 65.6 Imputed onlyb 0.45 -
 Urine microalbumin (mg/L)b 69.6 ± 160.7 73.0 ± 164.9 Imputed onlyb 0.21 -
 UACR (mg/g)b 407.9 ± 1,207.9 405.8 ± 1,187.0 29.0 ± 72.3 0.92 <0.001
 Sodium (mmol/L) 139.7 ± 3.3 139.8 ± 3.2 142.7 ± 2.4 (n = 366)a 0.71 <0.001
 Potassium (mmol/L) 4.6 ± 0.6 4.6 ± 0.6 4.6 ± 0.5 (n = 327)a 0.73 0.007
 Calcium (mg/dL) 9.1 ± 0.7 9.1 ± 0.7 9.7 ± 0.5 (n = 690)a 0.17 <0.001
 Phosphorus (mg/dL) 3.8 ± 1.0 3.8 ± 1.0 Not availablea 0.35 -
 AST (IU/L) 24.0 ± 28.0 24.4 ± 31.8 27.2 ± 18.6 0.25 <0.001
 ALT (IU/L) 21.6 ± 25.5 21.6 ± 29.6 23.2 ± 16.9 0.92 <0.001
 GGT (IU/L) 55.0 ± 100.7 56.0 ± 103.2 44.7 ± 110.3 (n = 712)a 0.54 0.011
 Serum bicarbonate/TCO2 (mEq/L) 25.4 ± 4.3 25.4 ± 4.3 Not availablea 0.25 -
 Total bilirubin (mg/dL) 0.68 ± 0.69 0.67 ± 0.59 0.56 ± 0.29 (n = 361)a 0.69 <0.001
 Total protein (g/dL) 7.00 ± 0.71 6.98 ± 0.73 7.35 ± 0.52 (n = 334)a 0.01 <0.001
Medical history
 RAAS inhibitor use 14,740 (74.6) 6,357 (75.1) 296 (7.2) 0.42 <0.001
 Statin use 14,056 (71.2) 6,008 (71.0) 439 (10.7) 0.75 <0.001
 Diagnosis of hypertensionc 14,740 (74.6) 6,357 (75.1) 1,565 (38.1) 0.42 <0.001
 Diabetes mellitus 1,708 (8.6) 739 (8.7) 1,055 (25.7) 0.84 <0.001
Additional comorbiditiesd
 Cardiovascular diseased 1,732 (6.1) 128 (3.2) - <0.001
 Cerebrovascular diseased 480 (1.7) 86 (2.2) - 0.103
 Chronic liver diseased 1,022 (3.6) Not available - -
 Malignancyd 3,126 (11.1) 110 (2.8) - <0.001
 Polycystic kidney diseased 1,487 (5.3) Not available - -
 Goutd 199 (0.7) 406 (10.3) - <0.001
 Subgroup: DKD 1,708/19,746 (8.6) 739/8,463 (8.7) Not available 0.84 -
 Subgroup: GN 637 (3.2) 316 (3.7) Not available 0.03 -
Outcomes
 Primary outcome (≥40% eGFR decline) 6,773 (34.3) 2,903 (34.3) 69 (1.7) >0.99  <0.001
 Secondary outcome (≥50% eGFR decline) 2,429 (12.3) 1,030 (12.2) 53 (1.3) 0.77 <0.001

Data are expressed as mean ± standard deviation or number (%).

ALT, alanine aminotransferase; AST, aspartate aminotransferase; DKD, diabetic kidney disease; eGFR, estimated glomerular filtration rate; GGT, gamma-glutamyl transferase; GN, glomerulonephritis; HbA1c, glycated hemoglobin; HDL, high-density lipoprotein; KoGES, Korean Genome and Epidemiology Study; RAAS, renin-angiotensin-aldosterone system; SNUH, Seoul National University Hospital; TCO2, total carbon dioxide; UACR, urine albumin-to-creatinine ratio.

aGGT, total bilirubin, total protein, serum sodium, and serum potassium were measured at baseline only; serum calcium was measured at baseline (n = 366) and the 4th follow-up survey (n = 439) only (total n = 690). Serum uric acid was measured only at the 4th, 6th, 7th, 8th, and 9th follow-up surveys (total n = 1,737). For all other waves, these variables were imputed using SNUH cohort medians (GGT, 27 IU/L; total bilirubin, 0.7 mg/dL; total protein, 7.2 g/dL; sodium, 139.7 mmol/L; potassium, 4.6 mmol/L; calcium, 9.2 mg/dL; uric acid, 5.8 mg/dL). Serum phosphorus and bicarbonate were not measured in any KoGES survey wave and are reported as “Not available.”

bUrine creatinine and urine microalbumin represent spot urine measurements included as independent model features alongside UACR. In the KoGES cohort, these variables were not available at any follow-up wave; all observations were set to SNUH cohort medians (urine creatinine, 102.8 mg/dL; urine microalbumin, 70.6 mg/L) for model inference. UACR was estimated from urine dipstick protein results using a semi-quantitative mapping: negative/trace = 15 mg/g; 1+ = 30 mg/g; 2+ = 100 mg/g; 3+ = 300 mg/g; 4+ = 1,000 mg/g.

cIn the SNUH derivation cohort, hypertension was defined using antihypertensive prescription records as a surrogate marker (RAAS inhibitor duration, >0 days; surrogate-based prevalence, 74.6%–75.1%). In the KoGES cohort, hypertension was defined based on physician diagnosis or current antihypertensive treatment history (prevalence, 38.1%), reflective of the community-based setting.

dAdditional comorbidities were ascertained from ICD-10 diagnosis records (SNUH, n = 28,209) or baseline self-reported medical history questionnaire matched via unique participant ID (KoGES, n = 3,960). These variables were not included as model input features. Chronic liver disease and polycystic kidney disease were not assessed in the KoGES questionnaire (not available).

Table 2.
Model performance in the internal and external validation cohorts
Model SNUH_AUROC (95% CI) SNUH_AUPRC KoGES_AUROC (95% CI) KoGES_AUPRC
XGBoost 0.937 (0.931–0.942) 0.898 0.863 (0.802–0.915) 0.473
LightGBM 0.937 (0.931–0.942) 0.899 0.856 (0.796–0.907) 0.469
CatBoost 0.939 (0.933–0.944) 0.899 0.846 (0.779–0.904) 0.460
RandomForest 0.936 (0.930–0.941) 0.893 0.867 (0.807–0.921) 0.535
Ensemble 0.939 (0.934–0.944) 0.901 0.859 (0.798–0.914) 0.489
KFRE 4-variable 0.879 (0.870–0.887) 0.832 0.882 (0.818–0.935) 0.545
KFRE 8-variable 0.884 (0.876–0.893) 0.837 0.889 (0.825–0.941) 0.550

Model performance in the internal and external validation cohorts. Comparison of discriminative performance metrics (AUROC and AUPRC) among individual ML models (XGBoost, LightGBM, CatBoost, Random Forest), the ensemble ML model, and the KFRE models. In the internal validation (SNUH), the ensemble ML model achieved the highest performance (AUROC, 0.939; 95% CI, 0.934–0.944), significantly outperforming the KFRE 8-variable model (AUROC, 0.884; 95% CI, 0.876–0.893). In the external validation (KoGES), the ensemble ML model maintained robust generalizability (AUROC, 0.859; 95% CI, 0.798–0.914), comparable to the KFRE models (4-variable AUROC, 0.882; 95% CI, 0.818–0.935). AUROC 95% CIs estimated by bootstrap resampling (1,000 iterations).

DeLong test comparing the ensemble ML model with KFRE models showed the following results: in SNUH, ensemble vs. KFRE 4-variable, ΔAUROC, +0.061 (95% CI, +0.055 to +0.067; p < 0.001), and ensemble vs. KFRE 8-variable, ΔAUROC, +0.057 (95% CI, +0.051 to +0.063; p < 0.001); in KoGES, ensemble vs. KFRE 4-variable, ΔAUROC, –0.023 (95% CI, –0.072 to +0.026; p = 0.36), and ensemble vs. KFRE 8-variable, ΔAUROC, –0.023 (95% CI, –0.072 to +0.026; p = 0.35) (Supplementary Table 3, available online).

AUPRC, area under the precision-recall curve; AUROC, area under the receiver operating characteristic curve; CI, confidence interval; KFRE, Kidney Failure Risk Equation; KoGES, Korean Genome and Epidemiology Study; ML, machine learning; SNUH, Seoul National University Hospital.

Table 3.
Baseline characteristics of clinical subgroups in the internal validation set
Variable DKD (n = 739) GN (n = 316) HTN (n = 660) Proteinuria-dominant (n = 1,063)
Demographics
 Age (yr) 65.4 ± 13.2 50.5 ± 16.8 68.0 ± 14.4 62.4 ± 15.9
 Body mass index (kg/m2) 24.5 ± 4.3 23.4 ± 3.7 24.4 ± 3.5 24.2 ± 4.5
 Male sex 467 (64.6) 140 (48.1) 353 (61.6) 633 (62.7)
Laboratory measurements
 eGFR (mL/min/1.73 m2) 38.6 ± 26.1 32.9 ± 31.7 40.2 ± 23.6 33.8 ± 24.9
 eGFR slope (mL/min/1.73 m2/yr) –2.1 ± 5.7 –0.9 ± 4.8 –1.7 ± 4.0 –2.6 ± 4.7
 Hemoglobin (g/dL) 11.7 ± 2.1 11.5 ± 2.2 12.2 ± 2.2 11.5 ± 2.3
 HbA1c (%) 7.1 ± 1.3 5.6 ± 0.7 6.3 ± 1.0 6.5 ± 1.3
 Serum albumin (g/dL) 4.0 ± 0.6 3.9 ± 0.5 4.1 ± 0.5 3.9 ± 0.6
 Glucose (mg/dL) 143.8 ± 65.0 101.2 ± 25.2 120.4 ± 51.6 127.7 ± 58.4
 Uric acid (mg/dL) 6.5 ± 2.0 6.8 ± 2.1 6.7 ± 2.0 6.8 ± 2.1
 Total cholesterol (mg/dL) 135.4 ± 50.6 166.4 ± 62.1 142.5 ± 48.3 150.8 ± 54.7
 Triglycerides (mg/dL) 149.6 ± 89.2 144.3 ± 90.7 141.4 ± 89.0 149.1 ± 87.5
 HDL cholesterol (mg/dL) 43.8 ± 13.7 53.5 ± 18.2 47.3 ± 14.7 46.8 ± 14.9
 Blood urea nitrogen (mg/dL) 35.8 ± 21.6 43.3 ± 28.0 32.6 ± 19.5 40.1 ± 24.8
 Serum creatinine (mg/dL) 3.2 ± 3.0 4.6 ± 4.3 2.6 ± 2.6 3.4 ± 2.9
 UACR (mg/g) 660.2 ± 1,813.9 623.5 ± 1,245.2 334.7 ± 989.8 1,753.1 ± 2,115.2
 Sodium (mmol/L) 139.5 ± 3.4 139.9 ± 3.0 140.0 ± 3.2 139.5 ± 3.4
 Potassium (mmol/L) 4.7 ± 0.6 4.7 ± 0.7 4.7 ± 0.6 4.7 ± 0.6
 Calcium (mg/dL) 9.1 ± 0.7 9.0 ± 0.7 9.1 ± 0.7 8.9 ± 0.7
 Phosphorus (mg/dL) 3.9 ± 1.1 4.3 ± 1.3 3.7 ± 0.9 4.0 ± 1.2
 AST (IU/L) 22.2 ± 11.0 21.8 ± 20.2 22.3 ± 11.2 22.8 ± 15.2
 ALT (IU/L) 21.0 ± 16.3 19.0 ± 20.7 20.7 ± 15.2 20.7 ± 21.9
 GGT (IU/L) 45.5 ± 80.5 34.6 ± 51.2 45.6 ± 66.6 58.6 ± 100.2
 Serum bicarbonate/TCO2 (mmol/L) 25.0 ± 4.4 24.0 ± 4.3 25.0 ± 4.2 24.4 ± 4.6
 Total bilirubin (mg/dL) 0.6 ± 0.3 0.6 ± 0.3 0.7 ± 0.3 0.6 ± 0.4
 Total protein (g/dL) 6.9 ± 0.7 6.8 ± 0.8 7.0 ± 0.6 6.8 ± 0.8
Medical history
 RAAS inhibitor use 617 (85.3) 275 (94.5) 526 (91.8) 912 (90.4)
 Statin use 592 (81.9) 244 (83.8) 465 (81.2) 830 (82.3)
 HTNa 140 (19.4) 275 (94.5) 660 (100) 68 (6.7)
 Diabetes mellitus 739 (100) 14 (4.8) 140 (24.4) 134 (13.3)
Outcomes
 Primary outcome (≥40% eGFR decline or ESRD) 303 (41.9) 163 (56.0) 171 (29.8) 524 (51.9)
 Secondary outcome (≥50% eGFR decline or ESRD) 111 (15.4) 54 (18.6) 60 (10.5) 215 (21.3)

Data are expressed as mean ± standard deviation or number (%). Groups are not mutually exclusive; patients may appear in multiple columns. Demographic and laboratory characteristics stratified by four clinically distinct subgroups: DKD, GN, HTN, and proteinuria-dominant disease (UACR ≥ 300 mg/g). The GN and proteinuria groups exhibited the most severe phenotypes, characterized by lower mean eGFR (32.9 and 33.8 mL/min/1.73 m2, respectively) and elevated serum creatinine, reflecting the high-risk nature of these etiologies in a tertiary care setting. Clinical subgroups were defined as follows based on the internal validation set: DKD, patients with a recorded diagnosis of diabetes mellitus (International Classification of Diseases, 10th Revision [ICD-10]: E10–E14; n = 739); GN, patients with a GN-specific ICD-10 subgroup code (N00–N08; n = 316); HTN, patients with an ICD-10 coded HTN diagnosis (I10–I15; n = 660); and proteinuria-dominant disease, patients with UACR ≥300 mg/g (n = 1,063).

ALT, alanine aminotransferase; AST, aspartate aminotransferase; DKD, diabetic kidney disease; eGFR, estimated glomerular filtration rate; ESRD, end-stage renal disease; GGT, gamma-glutamyl transferase; GN, glomerulonephritis; HbA1c, glycated hemoglobin; HDL, high-density lipoprotein; HTN, hypertension; RAAS, renin-angiotensin-aldosterone system; TCm2, total carbon dioxide; UACR, urine albumin-to-creatinine ratio.

aHTN was defined using antihypertensive prescription records as a surrogate marker.

Table 4.
Model performance by clinical subgroup from internal validation set
Subgroup Number Events (outcome ≥ 40%), n (%) Model AUROC (95% CI)
DKD 739 309 (41.8) Ensemble 0.946 (0.929–0.962)
KFRE 4-variable 0.861 (0.831–0.889)
KFRE 8-variable 0.866 (0.836–0.893)
GN 316 183 (57.9) Ensemble 0.941 (0.913–0.966)
KFRE 4-variable 0.908 (0.873–0.939)
KFRE 8-variable 0.912 (0.876–0.943)
HTN 660 203 (30.8) Ensemble 0.900 (0.870–0.928)
KFRE 4-variable 0.829 (0.793–0.867)
KFRE 8-variable 0.840 (0.804–0.876)
Proteinuria (UACR ≥300 mg/g) 1,063 533 (50.1) Ensemble 0.927 (0.912–0.941)
KFRE 4-variable 0.893 (0.874–0.912)
KFRE 8-variable 0.895 (0.876–0.915)

AUROC, area under the receiver operating characteristic curve; CI, confidence interval; DKD, diabetes mellitus; GN, glomerulonephritis; HTN, hypertension; KFRE, Kidney Failure Risk Equation; UACR, urine albumin-to-creatinine ratio.

95% CIs were estimated by bootstrap resampling (1,000 iterations). Ensemble denotes the soft-voting ensemble of XGBoost, LightGBM, CatBoost, and Random Forest. Subgroups are not mutually exclusive and were defined as follows: DKD, International Classification of Diseases, 10th Revision [ICD-10]: E10–E14; GN, ICD-10: N00–N08; HTN, ICD-10: I10–I15; proteinuria-dominant, UACR ≥300 mg/g.

References

1. Tangri N, Stevens LA, Griffith J, et al. A predictive model for progression of chronic kidney disease to kidney failure. JAMA 2011;305:1553–1559.
crossref pmid
2. Takkavatakarn K, Oh W, Cheng E, Nadkarni GN, Chan L. Machine learning models to predict end-stage kidney disease in chronic kidney disease stage 4. BMC Nephrol 2023;24:376.
crossref pmid pmc pdf
3. Elshewey AM, Selem E, Abed AH. Improved CKD classification based on explainable artificial intelligence with extra trees and BBFS. Sci Rep 2025;15:17861.
crossref pmid pmc pdf
4. Li C, Liu J, Fu P, Zou J. Artificial intelligence models in diagnosis and treatment of kidney diseases: current status and prospects. Kidney Dis (Basel) 2025;11:491–507.
crossref pmid pmc pdf
5. Wainstein M, Rahimi AK, Katz I, et al. A comparison between a random forest model and the kidney failure risk equation to predict progression to kidney failure. medRxiv [Preprint] 2023 May 17 [cited 2026 Feb 2]. Available from: https://doi.org/10.1101/2023.05.16.23290068
crossref
6. Moosavi Kashani S, Zargar Balaye Jame S. Comparing the performance of machine learning models in predicting the risk of chronic kidney disease. J Arch Mil Med 2023;11:e140885.
7. Miller ZA, Dwyer K. Artificial intelligence to predict chronic kidney disease progression to kidney failure: a narrative review. Nephrology (Carlton) 2025;30:e14424.
crossref pmid
8. Leung KC, Ng WW, Siu YP, Hau AK, Lee HK. Deep learning algorithms for predicting renal replacement therapy initiation in CKD patients: a retrospective cohort study. BMC Nephrol 2024;25:95.
crossref pmid pmc pdf
9. Barbour SJ, Coppo R, Zhang H, et al. Evaluating a new international risk-prediction tool in IgA nephropathy. JAMA Intern Med 2019;179:942–952.
crossref pmid pmc
10. Kim Y, Han BG, KoGES group. Cohort profile: the Korean Genome and Epidemiology Study (KoGES) Consortium. Int J Epidemiol 2017;46:e20.
crossref pmid pmc
11. Bai Q, Su C, Tang W, Li Y. Machine learning to predict end stage kidney disease in chronic kidney disease. Sci Rep 2022;12:8377.
crossref pmid pmc pdf
12. Levey AS, Gansevoort RT, Coresh J, et al. Change in albuminuria and GFR as end points for clinical trials in early stages of CKD: a scientific workshop sponsored by the National Kidney Foundation in collaboration with the US Food and Drug Administration and European Medicines Agency. Am J Kidney Dis 2020;75:84–104.
crossref pmid
13. Xie X, Liu Y, Perkovic V, et al. Renin-angiotensin system inhibitors and kidney and cardiovascular outcomes in patients with CKD: a Bayesian network meta-analysis of randomized clinical trials. Am J Kidney Dis 2016;67:728–741.
crossref pmid
14. Li Y, Al-Sayouri S, Padman R. Towards interpretable end-stage renal disease (ESRD) prediction: utilizing administrative claims data with explainable AI techniques. AMIA Annu Symp Proc 2025;2024:664–673.
pmid pmc
15. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 2006;26:565–574.
crossref pmid pmc pdf
16. Park RW. The distributed research network, observational health data sciences and informatics, and the South Korean Research Network. Korean J Med 2019;94:309–314.
crossref pdf
17. Voss EA, Makadia R, Matcho A, et al. Feasibility and utility of applications of the common data model to multiple, disparate observational health databases. J Am Med Inform Assoc 2015;22:553–564.
crossref pmid pmc pdf
18. Hripcsak G, Duke JD, Shah NH, et al. Observational Health Data Sciences and Informatics (OHDSI): opportunities for observational researchers. Stud Health Technol Inform 2015;216:574–578.
pmid pmc
19. Kim JW, Kim C, Kim KH, et al. Scalable infrastructure supporting reproducible nationwide healthcare data analysis toward FAIR stewardship. Sci Data 2023;10:674.
crossref pmid pmc pdf
20. Cook NR. Use and misuse of the receiver operating characteristic curve in risk prediction. Circulation 2007;115:928–935.
crossref pmid pmc
21. Gomez-Cabello CA, Borna S, Pressman S, Haider SA, Haider CR, Forte AJ. Artificial-intelligence-based clinical decision support systems in primary care: a scoping review of current clinical implementations. Eur J Investig Health Psychol Educ 2024;14:685–698.
crossref pmid pmc


ABOUT
BROWSE ARTICLES
EDITORIAL POLICY
FOR CONTRIBUTORS
Editorial Office
#301, (Miseung Bldg.) 23, Apgujenog-ro 30-gil, Gangnam-gu, Seoul 06022, Korea
Tel: +82-2-3486-8736    Fax: +82-2-3486-8737    E-mail: registry@ksn.or.kr                

Copyright © 2026 by The Korean Society of Nephrology.

Developed in M2PI

Close layer