Kidney Res Clin Pract > Epub ahead of print
Ye, Zhang, Xia, Yang, Luo, Tang, He, Huang, Xiao, and Ren: An interpretable multimodal fusion model for improving diagnosis and prognosis prediction of kidney allograft rejection

Abstract

Background

Kidney allograft rejection exhibits highly complex features due to the intricate immune mechanisms involved. The early and accurate diagnosis of rejection remains challenging, particularly when differentiating between/among subtypes. Herein, the authors present “RenalTransPredNet,” an interpretable model that integrates whole-slide images with routine clinicopathological variables to diagnose kidney allograft rejection.

Methods

This study included kidney allograft biopsies performed at the Headquarters and Lingnan Hospital of the Third Affiliated Hospital of Sun Yat-sen University between January 2015 and August 2024. After screening, 1086 digital whole-slide images from 362 biopsies and their corresponding clinicopathological information were analyzed, comprising 209 other lesions and 153 rejections. By integrating multiple instance learning-based pathological image analysis with random forest-based clinicopathological information analysis, RenalTransPredNet was used to efficiently construct models for kidney allograft rejection-related multiclass and binary classification tasks.

Results

RenalTransPredNet achieved a higher area under the curve (AUC) value than both the pathological imaging and clinicopathological information models in detecting and subtyping kidney allograft rejection, yielding an AUC of 0.886 using data from the internal testing set. For the prediction of treatment response to rejection, RenalTransPredNet yielded an AUC of 0.743. In the task predicting graft loss after rejection, RenalTransPredNet demonstrated predictive capabilities for graft loss at 1, 2, 3, and 5 years post-rejection, with AUCs of 0.876, 0.942, 0.904, and 0.978, respectively.

Conclusion

These findings support the potential of RenalTransPredNet in improving the diagnostic accuracy of pathologists in detecting and subtyping kidney allograft rejection and in contributing to prognostic assessments.

Introduction

Chronic kidney disease affects approximately 10% of the global population, with millions progressing to end-stage renal disease (ESRD) [1,2]. Kidney transplantation remains the preferred treatment for ESRD [3]. However, posttransplant intrinsic kidney injury poses significant threats to graft survival, with rejection being the primary concern [4]. Notably, kidney allograft rejection accounts for 63% of graft loss cases occurring >1 year posttransplantation [5]. Therefore, an early and accurate diagnosis of rejection may optimize posttransplant management and prolong graft survival.
Histopathological assessment of allograft biopsies remains the cornerstone modality for diagnosing rejection [6]. This process involves detecting and subtyping rejection as T cell-mediated rejection (TCMR), antibody-mediated rejection (ABMR), or mixed TCMR/ABMR (MIXR) [7]. The prognosis of rejection is further stratified according to lesion severity. The subtyping and prognosis of rejection guide therapeutic decision-making and determine clinical management. However, manual assessment performed by transplant pathologists is inherently subjective, with inter-observer variability limiting reproducibility [8]. Moreover, this process is time-consuming, expensive, and labor-intensive. Artificial intelligence (AI)-assisted digital pathology systems may overcome these limitations by providing more objective, efficient, and standardized analyses, and improving accuracy, consistency, and workflow.
Multiple instance learning (MIL) is a weakly supervised machine-learning paradigm that is well-suited for digital histopathological image analysis [9,10]. MIL effectively extracts discriminative features from whole-slide images (WSIs) using only slide-level labels, maintaining high diagnostic accuracy while significantly reducing the annotation burden [11]. Recently, we developed MIL models based on WSIs for automated assessment of transplant pathology, specifically for kidney allograft rejection [12]. However, these models can be further improved. First, previously proposed MIL models did not incorporate readily accessible clinicopathological variables that could aid in the diagnosis of rejection, including serum biomarkers such as creatinine (CR). Second, clinical decision-making is inherently a multimodal integration process [13]. However, to the best of our knowledge, no previous study has systematically integrated multimodal information, including pathological images and clinicopathological texts, into a risk stratification framework for renal allograft rejection.
As such, the present study aimed to develop and validate “RenalTransPredNet,” an integrative multimodal information system, for the diagnosis of kidney allograft rejection. RenalTransPredNet combines WSIs with routine clinicopathological parameters to enhance predictive accuracy for the detection, subtyping, and prognosis of kidney allograft rejection. In addition to classifying lesions and predicting rejection prognosis, RenalTransPredNet provides a visual interpretation of the impact of each parameter on the predictions using Shapley additive explanation (SHAP) plots.

Methods

Study design and participants

The present study was approved by the Institutional Review Board of the Third Affiliated Hospital of Sun Yat-sen University in Guangdong, China (No. 2023-041-01). Requirements for written informed consent were waived due to the retrospective design of the study and the use of anonymized data.
Patients were eligible if they had undergone kidney transplant biopsies at the Headquarters and Lingnan Hospital of the Third Affiliated Hospital of Sun Yat-sen University (SYSUTH) between January 2015 and August 2024. There is no fixed schedule for biopsy, which is primarily performed when clinically indicated, such as in cases of elevated serum CR or proteinuria. The exclusion criteria were as follows: kidney transplantation before 2015; time-zero biopsy; and repeat biopsy.

Data collection and data set curation

Hematoxylin and eosin (H&E) stained slides were digitized at 40× original magnification using a digital slide scanner (Panoramic 250 FLASH; 3DHISTECH Ltd.). Slides exhibiting insufficient renal cortical regions, poor focalization, or degraded staining were excluded. Clinical and pathological data were retrieved from the institutional electronic medical records system. Information from 1,086 H&E-stained WSIs and routine clinicopathological parameters from 362 patients was included in the dataset. The patient enrollment process is illustrated in Fig. 1. Patients were randomly allocated to the training and independent testing sets at a 7:3 ratio. The compositions of the training and testing sets of the models are listed in Supplementary Table 1 (available online).
Two experienced transplant pathologists individually reviewed the biopsy samples and classified them into the following categories based on consensus: TCMR, ABMR, MIXR, and other lesions (“others”). In cases of discrepant evaluations between the two pathologists, a third pathologist was consulted to reach a consensus. All diagnoses were made in accordance with the Banff 2022 criteria [14], and the diagnostic details are summarized in Supplementary Table 2 (available online). Diagnostic confirmation and region of interest annotation were independently performed by an additional subspecialist using an automated slide analysis platform (ASAP 1.9; Radboud University Medical Center) for WSI analysis. The region of interest annotation process is shown in Supplementary Fig.  1 (available online). Diagnoses of all cases are reported in Supplementary Table 3 (available online).
The prognosis of rejection was evaluated based on the treatment response and kidney allograft loss. Graft loss was defined as resumption of chronic dialysis [15]. The treatment response for rejection was determined by the return of the estimated glomerular filtration rate (eGFR) to within 10% of the baseline level at 3 months post-biopsy [16]. The specific treatment protocol for patients with rejection is detailed in the Supplementary Methods (available online). The variation ratio of serum CR was defined as the percentage increase in CR at biopsy compared with the baseline level. Baseline eGFR or CR was assessed within 3 months preceding biopsy. C4d positivity was determined by linear staining of >0 peritubular capillaries using immunohistochemistry on formalin-fixed paraffin-embedded tissues.

Model development

Pathological image-based multiple instance learning model

Pathological image analysis was performed using an MIL framework in which image features were extracted using a pretrained ResNet50 model. MIL represents a methodology for processing data bags containing multiple instances, and is particularly suitable for pathological image analysis because each WSI folder may encompass numerous image instances that collectively reflect a patient’s pathological status [17]. In this study, each WSI folder was considered to be a bag containing multiple image instances. For individual instances, the features extracted using the ResNet50 model were flattened and aggregated to form bag-level feature representations. Subsequently, these features underwent global average pooling to generate the final bag features for the downstream classification tasks.
More specifically, the ResNet50 model was used as the foundational feature extractor, and preloaded with ImageNet weights, and modified by removing the top layers to adapt to the feature extraction requirements. For each image, the tensorflow.keras.preprocessing.image module was initially used to load and resize the images to 224 × 224 pixels, followed by preprocessing to fulfill the ResNet50 input specifications. After feature extraction, the multidimensional features were flattened into one-dimensional arrays. For each bag, average pooling was applied across all instance features to derive bag-level representations. This process was implemented using a self-defining function, create_instance_bags, which transforms the bag data into model-acceptable formats. The model was designed to accept bag-level feature representations as inputs and generate probability distributions across classes as outputs. The architecture incorporated a GlobalAveragePooling2D layer and fully connected dense layers. The training used the Adam optimizer with a learning rate of 0.001 using categorical cross-entropy as the loss function and accuracy as the evaluation metric.

Clinicopathological information-based random forest model

For clinicopathological text data, the following five key clinicopathological indicators were standardized using the standard scaler: age, donor-specific antibody (DSA) status, C4d staining status, variation ratios of serum CR, and lymphocyte percentage (LYMPH). These features were individually extracted and standardized using StandardScaler, which transforms the values of each feature into a distribution with a mean of 0 and standard deviation of 1. The standardized features were stored in a new Pandas DataFrame for subsequent modeling. Simultaneously, the target variables (pathological subtype, treatment response, and graft loss) were extracted from the original dataset to serve as labels for supervised learning.
Analysis of clinicopathological information was performed using a random forest algorithm to model standardized clinicopathological features. Random forest, an ensemble learning method based on decision trees, improves classification stability and accuracy by constructing multiple decision trees and aggregating their predictions, making it particularly effective for managing high-dimensional data [18,19]. In this study, a RandomForestClassifier model trained on standardized clinical features was implemented to predict post-renal transplantation outcomes. More specifically, the preprocessed clinicopathological dataset was partitioned into training (70%) and test (30%) sets. The random forest model was trained on the training set, and its classification performance was evaluated using an independent test set. Model accuracy was calculated as the proportion of correctly predicted samples relative to the total sample size. To elucidate the decision-making process and identify the clinicopathological indicators contributing to predictions, a feature importance analysis was performed using SHAP. Rooted in Shapley values, SHAP quantifies the individual feature contributions to model predictions, thereby enhancing interpretability. The trained random forest model was interpreted using Shap.TreeExplainer, with the SHAP values computed for each sample. A summary plot of feature importance (shap.summary_plot) was subsequently generated to visually delineate the magnitude of the impact of each clinicopathological feature on the model predictions.

RenalTransPredNet: Integrated model construction

RenalTransPredNet is an integrated model that enhances predictive performance by synergistically combining outcomes from pathological image analysis with clinicopathological information processing. Specifically, this architecture generates final predictions through the probability averaging of outputs from both an MIL model and a random forest model. This strategy not only capitalizes on the complementary nature of histopathological imaging and clinicopathological data but also improves model generalizability via multimodality data integration.
During the construction of the ensemble model, the MIL-based pathological image analysis model and the random forest-driven clinicopathological data analysis model were initially trained separately. The predictive probabilities from the test-set performance of each model were extracted systematically. Subsequently, probabilistic fusion was achieved through arithmetic averaging of the outputs of both models, which is a computationally efficient approach designed to leverage cross-modal advantages while mitigating the limitations of individual modalities. The construction process of RenalTransPredNet is illustrated in Fig. 2.
RenalTransPredNet was designed based on the hypothesis that histopathological images and clinicopathological parameters represent the pathological state and clinical manifestations of patients with rejection from different perspectives. Pathological slides provide tissue-level morphological details, whereas clinical metrics capture systemic physiological status and therapeutic responses. Through multimodal fusion, RenalTransPredNet provides an integrated assessment of kidney allograft rejection, contributing to improved diagnostic performance.

Prediction tasks and model application

RenalTransPredNet was designed to address three critical predictive tasks in kidney transplantation management: four-category histopathological lesion classification, integrating histopathological images and clinicopathological parameters to classify posttransplant pathological lesions into four subtypes (Other, TCMR, ABMR, and MIXR), thereby achieving the detection and subtyping of rejection and treatment response prediction, developing a therapeutic response model in patients with rejection to guide clinical decision making; and long-term graft survival prediction, time-specific predictive models for graft loss within 1-, 2-, 3-, and 5-year post-rejection intervals were developed to aid in clinical management.

Statistical analysis

The performance of RenalTransPredNet was evaluated using multiple metrics. The area under the curve (AUC) was computed for all tasks to assess the predictive capability of the model. In all tasks, a confusion matrix was calculated to analyze the performance of the model in each category, reflecting its predictive accuracy for individual classes. For all classification tasks, a detailed comparison of the accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1-score was performed.

Results

Baseline information

There were 1,086 WSIs of 209 other lesions (57.7%) and 153 rejections (42.3%) from the 362 biopsies (Table 1). Each biopsy specimen had three WSIs corresponding to the three layers of the tissue block. The median ages of patients with Other, TCMR, ABMR, and MIXR were 39, 36, 45, and 44 years, respectively. The median variation ratios of serum CR were 0.12, 0.26, 0.23, and 0.45, respectively. The median LYMPHs were 0.18, 0.18, 0.23, and 0.16, respectively. The positivity rates for DSA were 1%, 0%, 16%, and 29%, respectively. The positivity rates for C4d were 13%, 0%, 76%, and 56%, respectively.

Performance of pathological image and clinicopathological information models in classifying histopathological lesions

In the internal testing set, the pathological image model yielded an AUC of 0.831 for the classification of the four-category renal allograft lesions, whereas the clinicopathological information model yielded an AUC of 0.801. Confusion matrices for both models are presented in Fig. 3A and B, showing the results of the classification of the four-category renal allograft lesions. The diagonal entries indicate the proportion of correct predictions, whereas the off-diagonal entries reflect misclassification rates. The deeper the blue shading, the higher the diagnostic accuracy.

Performance of RenalTransPredNet in classifying histopathological lesions

As shown in Fig. 3, with the integration of image-based MIL and clinicopathological information predictions, RenalTransPredNet demonstrated a higher performance (AUC, 0.886) than the pathological image model (AUC, 0.831) and clinicopathological information model (AUC, 0.801) for the overall four-category classification in the internal testing set. Furthermore, RenalTransPredNet demonstrated higher performance across all other metrics (Supplementary Table 4, available online). The variation ratios of the serum CR predictions weighed most importantly for the quadruple-classification decision prediction of RenalTransPredNet among the clinicopathological parameters, followed by C4d, patient age, LYMPH, and DSA (Fig. 4).
RenalTransPredNet misclassified 15 rejection cases into Other. An in-depth analysis of the 15 false-negative cases (seven TCMR and eight ABMR) revealed that 12 (80.0%) exhibited mild pathological grades (e.g., borderline or grade IA TCMR or early-stage ABMR with limited microvascular inflammation). In 10 out of the 15 cases, the CR levels were relatively stable (variation <15%), providing weak functional signals for the fusion model. These suggest a diagnostic challenge when both morphological features and clinical markers are subtle.

Performance of RenalTransPredNet (whole-slide image-only) in classifying histopathological lesions in the stable creatinine group

To evaluate whether our model offers independent discriminating value beyond routine clinical markers, we conducted a sensitivity analysis in the subgroup with stable renal function. Even within CR variation of less than 20% relative to baseline, RenalTransPredNet (WSI-only) achieved a high AUC of 0.817 in identifying rejection cases (Supplementary Fig. 2, available online).

Performance of RenalTransPredNet in predicting treatment response

Based on the internal testing set, Fig. 5A shows the confusion matrices of RenalTransPredNet for predicting the treatment response of patients with rejection. RenalTransPredNet yielded an AUC of 0.743 in predicting treatment response (Fig. 5B), with additional performance metrics provided in Table 2. As shown in Fig. 6, serum CR variation ratios exhibited the highest predictive influence on the treatment response prediction decisions of RenalTransPredNet among the clinicopathological parameters, with subsequent contributions from LYMPH, patient age, C4d, and DSA.

Performance of RenalTransPredNet in predicting graft loss

RenalTransPredNet was used to predict graft loss at 1, 2, 3, and 5 years after rejection. The confusion matrices of the prediction models are presented in Fig. 7A. In the internal testing set, RenalTransPredNet demonstrated favorable prediction capabilities for graft loss at 1, 2, 3, and 5 years after rejection (Fig. 7B), with AUCs of 0.876, 0.942, 0.904, and 0.978 and accuracies of 0.92, 0.97, 0.84, and 0.91, respectively (Table 2). The serum CR variation ratio predictions were the most important among the clinicopathological parameters regarding the graft loss prediction of RenalTransPredNet (Fig. 8), followed by LYMPH, patient age, C4d, and DSA.

Discussion

Kidney transplant rejection exhibits highly complex features due to the intricate immune mechanism involved [20]. Diagnosing rejection based solely on histopathological images remains challenging because differential diagnosis requires the integration of multimodal information. The diagnostic utility of image-based MIL predictions and clinical parameters has previously been explored for the diagnosis of kidney allograft rejection. However, the integration of multimodal features for improved diagnosis remains poorly understood. In this study, we introduced RenalTransPredNet, which integrates image-based MIL prediction and clinicopathological parameters to predict renal allograft rejection. For classifying kidney allograft lesions into four categories, RenalTransPredNet achieved higher AUC values than either single-modality approach, with an AUC of 0.886 in the internal testing set. Moreover, the model demonstrated predictive performance for treatment response in patients with rejection, as well as for assessing post-rejection graft loss across multiple time points. These results support the potential of RenalTransPredNet as an AI-assisted pathology analysis tool that enhances pathologists' accuracy in distinguishing between renal rejection and other allograft lesions, as well as in assessing rejection prognosis.
The strength of this study lies in the integration of image-based MIL prediction with routine clinicopathological parameter prediction to ensure complementary information and comprehensive feature representation. By classifying renal allograft lesions into four categories (i.e., TCMR/ABMR/MIXR/Other), RenalTransPredNet enabled the detection and subtyping of rejection and achieved a higher AUC value than the pathological image and clinicopathological information models in the independent internal testing set (AUC, 0.886 versus 0.831 or 0.801). Kers et al. [21] constructed deep-learning models using WSIs to detect kidney rejection and reported that their model achieved an AUC of 0.78 in a multicenter setting. Ye et al. [12] developed a histological image-based MIL model to classify renal rejection subtypes and achieved an AUC of 0.798 in a single-center setting. However, these studies relied solely on pathological images, which limited the comprehensive characterization of rejection features. We constructed RenalTransPredNet by integrating rejection characteristics across different dimensions, resulting in improved performance. These results confirm the feasibility of multimodal fusion for enhancing the model performance.
Additionally, we evaluated RenalTransPredNet for prognostic prediction in patients with rejection. In the internal testing sets, RenalTransPredNet achieved a high performance in terms of treatment response and graft loss. Both indicators were closely correlated with allograft prognosis. These indicators reflect kidney allograft function by assessing eGFR variations and dialysis resumption [22,23]. Here, we found that RenalTransPredNet demonstrated favorable performance for the 5-year survival of renal allografts after rejection (Table 2). These results suggest that RenalTransPredNet may have the potential to contribute to prognostic assessments of rejection. Unfortunately, long-term follow-up data remain unavailable because the study was limited to cases enrolled over the previous decade.
Interpretability optimization was implemented using RenalTransPredNet. In our previous study, heat maps highlighted the key discriminative features of the pathological image model, including interstitial inflammation and tubulitis in TCMR, and peritubular capillaritis in ABMR [12]. To identify key clinicopathological parameters, our statistical analysis of variables across different rejection subtypes and other lesions revealed significant differences in the following: age, CR variation ratio, months from transplantation to biopsy, C4d, DSA, and LYMPH. Although months from transplantation to biopsy differed significantly, correlation analysis revealed a significant association between months from transplantation to biopsy and chronicity scores (e.g., interstitial fibrosis [ci] and tubular atrophy [ct]) derived from WSIs (Supplementary Fig. 3, available online). Thus, to avoid multicollinearity and improve the biological plausibility of predictions, our fusion model prioritized WSIs’ features that directly reflect tissue injury, while excluding the variable of months from transplantation to biopsy. The SHAP plots further show the relative contribution of each clinicopathological parameter to the RenalTransPredNet predictions. We observed that the CR variation ratio had the greatest effect on the decisions made by RenalTransPredNet for all clinicopathological parameters. This parameter reflects alterations in renal allograft function and serves as an indicator of rejection. Interestingly, it also plays a significant role in predicting the treatment response and renal allograft survival. C4d and LYMPH also made significant contributions in different predictive tasks. DSA and patient age had less of an impact on the decision. C4d and DSA are key elements in the pathological diagnosis of ABMR. However, DSA positivity was detected in 5.2% of patients (19/362) who underwent kidney allograft biopsies, which may partially account for the relatively low contribution of DSA. LYMPH reflects the patient’s immune status to some extent, which may influence their response to treatment and graft survival.
To evaluate whether our model provides independent diagnostic information beyond routine clinical markers, we stratified patients by renal function stability. Notably, even in patients with stable CR (variation <20%), the WSI-only model maintained a high discriminative ability (AUC, 0.817) and successfully identified subclinical rejection cases that would have been missed by CR monitoring alone. This suggests that WSI-derived morphological features can detect early rejection prior to significant functional deterioration, offering a window for timely intervention. Although further validation in larger cohorts is needed, these results provide evidence that our model adds independent and incremental value to conventional monitoring, potentially addressing an unmet need in non-invasive surveillance of transplant recipients.
The present study had some limitations. First, although RenalTransPredNet demonstrated high diagnostic performance across diverse clinical tasks, its generalizability is limited due to the lack of external validation. Additionally, as RenalTransPredNet was developed primarily using a cohort of clinical indication biopsies, its applicability to general screening or protocol biopsies requires further validation. Second, even in RenalTransPredNet, some cases with rejection were misclassified into other categories. As shown in the results, these misclassified cases shared common features: subtle morphological changes and atypical clinical presentations. This indicates that while RenalTransPredNet is highly sensitive to typical rejection, its sensitivity to minimal or subclinical lesions remains a limitation. Future work should include external multicenter validation and enlarge the training set with mild/borderline rejection cases to improve discrimination of subtle rejection phenotypes, including early ABMR. Third, other parameters, including donor characteristics, human leukocyte antigen mismatch counts, and immunosuppressive regimens, are important for clinical diagnostic tasks. Regrettably, retrospective studies have difficulty in collecting detailed information. The utilities of these parameters will be explored further in future prospective studies. Fourth, imaging examinations, such as ultrasound and magnetic resonance imaging, also play important roles in the diagnosis of kidney rejection, and the integration of these multimodal images may further improve the performance of RenalTransPredNet.
This study demonstrated the utility of RenalTransPredNet for the detection, subtyping, and prognostic prediction of kidney allograft rejection. RenalTransPredNet integrated pathological images with clinicopathological information and showed higher performance with an internal dataset, supporting the importance of multimodal integration in the diagnosis of kidney allograft rejection. The SHAP analysis further presented the interpretation of the decision of RenalTransPredNet and demonstrated the importance of each parameter. These findings suggest that RenalTransPredNet may have the potential to assist in improving pathologist detection and subtyping accuracy for renal allograft rejection and to contribute to prognostic assessments. Future prospective studies are needed to validate its clinical applicability.

Notes

Conflicts of interest

All authors have no conflicts of interest to declare.

Funding

This work was supported by theShenzhen Science and Technology Program (No. JCYJ20220530145001002) and the Shenzhen Medical Research Fund (No. C2401018).

Data sharing statement

All data are available in the manuscript and Supplementary Materials. Further inquiries can be directed to the corresponding author.

Authors’ contributions

Conceptualization: YR, HX, ZH (Huang), YY

Data curation: YY, LZ, LX, YL, ZT

Formal analysis: YR, SY, ZH (He)

Funding acquisition: YR

Project administration: YY, LZ, LX

Writing–original draft: YY

Writing–review & editing: YR, HX, ZH (Huang)

All authors read and approved the final manuscript.

Figure 1.

The flowchart of study design and data curation.

Kidney allograft biopsies from two independent hospitals (Headquarters and Lingnan Hospital of SYSUTH) were included in this study. After exclusion, 1,086 whole-slide images (WSIs) together with their clinicopathological parameters from 362 digital kidney allograft biopsies were used in the analysis.
ABMR, antibody-mediated rejection; KT, kidney transplantation; SYSUTH, Third Affiliated Hospital of Sun Yat-sen University; TCMR, T cell-mediated rejection.
j-krcp-25-397f1.jpg
Figure 2.

The construction of RenalTransPredNet and its application in multiple clinical tasks.

Pathological image analysis employs a multiple instance learning framework, and clinicopathological information (CI) analysis utilizes a random forest algorithm. The multimodal model integrates both analyses and is applied in the diagnosis of rejection and the prediction of the prognosis.
ABMR, antibody-mediated rejection; CR, creatinine; DSA, donor-specific antibody; IHC, immunohistochemistry; LYMPH, lymphocyte percentage; MIXR, mixed TCMR/ABMR; SHAP, Shapley additive explanation; TCMR, T cell-mediated rejection; WSI, whole-slide image.
j-krcp-25-397f2.jpg
Figure 3.

The performance of the three models in classifying four-category renal allograft lesions.

(A–C) Classification results are shown by confusion matrices of the three models in the independent internal testing set. The diagonal entries indicate the proportion of correct predictions, whereas the off-diagonal entries reflect misclassification rates. The deeper the blue shading, the higher the diagnostic accuracy. (D) The receiver operator characteristic (ROC) curves and area under the curve (AUC) value of the models of the pathological image model (red line), clinicopathological information model (orange line), and RenalTransPredNet (blue line) in the independent internal testing set. RenalTransPredNet performed the best, with an AUC value of 0.886.
ABMR, antibody-mediated rejection; MIXR, mixed TCMR/ABMR; TCMR, T cell-mediated rejection.
j-krcp-25-397f3.jpg
Figure 4.

Global Shapley values for the interpretation of RenalTransPredNet in classifying four-category renal allograft lesions.

CR, creatinine; DSA, donor-specific antibody; LYMPH, lymphocyte percentage; SHAP, Shapley additive explanation.
j-krcp-25-397f4.jpg
Figure 5.

The performance of RenalTransPredNet in predicting the treatment response of rejection.

(A) The classification result is shown by the confusion matrix of RenalTransPredNet in the independent internal testing set. (B) The receiver operator characteristic (ROC) curve of RenalTransPredNet in the internal testing set, with an area under the curve (AUC) value of 0.743.
j-krcp-25-397f5.jpg
Figure 6.

Global Shapley values for the interpretation of RenalTransPredNet in predicting treatment response of rejection.

CR, creatinine; DSA, donor-specific antibody; LYMPH, lymphocyte percentage; SHAP, Shapley additive explanation.
j-krcp-25-397f6.jpg
Figure 7.

The performance of RenalTransPredNet in predicting graft loss post-rejection.

(A) Classification results are shown by confusion matrices of RenalTransPredNet in the independent internal testing set. (B) The receiver operator characteristic (ROC) curves in the internal testing set for RenalTransPredNet to predict graft loss at 1 (red line), 2 (orange line), 3 (blue line), and 5 years (green line) post-rejection, with area under the curves (AUCs) of 0.876, 0.942, 0.904, and 0.978, respectively.
j-krcp-25-397f7.jpg
Figure 8.

Global Shapley values for the interpretation of RenalTransPredNet in predicting graft loss post-rejection.

CR, creatinine; DSA, donor-specific antibody; LYMPH, lymphocyte percentage; SHAP, Shapley additive explanation.
j-krcp-25-397f8.jpg
Table 1.
Distribution of the patient characteristics
Characteristic Other lesions Rejection (n = 153)
TCMR ABMR Mixed rejections
No. of patients 209 64 62 27
Male sex 153 (73.2) 43 (67.2) 46 (74.2) 17 (63.0)
Age (yr) 39 (32.5–48.0) 36 (31.0–44.8) 45 (35.8–56.0) 44 (38.0–55.0)
Time from transplantation to biopsy (mo) 18.0 (5.5–49.8) 7.0 (3.7–17.7) 47.9 (11–72.3) 22.1 (6.6–50.1)
Follow-up after biopsy (mo) 17.6 (7.2–58.4) 16.4 (3.4–50.1) 12.9 (6.0–20.0) 1.5 (0–17.0)
Variations of serum CR 0.12 (0.07–0.21) 0.26 (0.14–0.66) 0.23 (0.12–0.36) 0.45 (0.22–1.09)
LYMPH 0.18 (0.12–0.25) 0.18 (0.11–0.24) 0.23 (0.17–0.28) 0.16 (0.14–0.24)
Renal artery RI 0.73 (0.68–0.79) 0.75 (0.7–0.78) 0.74 (0.69–0.79) 0.78 (0.67–0.86)
Urine protein positive 96 (45.9) 25 (39.1) 36 (58.1) 15 (55.6)
Graft loss 85 (40.7) 24 (37.5) 21 (33.9) 20 (74.1)
Treatment response 32 (50.0) 29 (47.0) 4 (15.0)
DSA
 Absent 199 (95.2) 59 (92.2) 37 (59.7) 11 (40.7)
 HLA I class 0 (0) 0 (0) 2 (3.2) 0 (0)
 HLA II class 3 (1.4) 0 (0) 6 (9.7) 6 (22.2)
 HLA I & II class 0 (0) 0 (0) 2 (3.2) 2 (7.4)
 No available 7 (3.3) 5 (7.8) 15 (24.2) 8 (29.6)
Preformed DSA 0 (0) 0 (0) 2 (3.2) 0 (0)
De novo DSA 3 (1.4) 0 (0) 8 (12.9) 8 (29.6)
C4d score
 0 181 (86.6) 64 (100) 15 (24.2) 12 (44.4)
 1 26 (12.4) 0 (0) 24 (38.7) 4 (14.8)
 2 2 (1.0) 0 (0) 15 (24.2) 3 (11.1)
 3 0 (0) 0 (0) 8 (12.9) 8 (29.6)

Data are expressed as number only, number (%), or median (interquartile range).

ABMR, antibody-mediated rejection; CR, creatinine; DSA, donor-specific antibody; HLA, human leukocyte antigen; LYMPH, lymphocyte percentage; RI, resistive index; TCMR, T cell-mediated rejection.

Table 2.
Performance of RenalTransPredNet in binary classification tasks
RenalTransPredNet AUC ACC PPV NPV SENS SPEC F1
Treatment response 0.74 0.77 0.79 0.76 0.61 0.88 0.77
Graft loss, 1 year 0.88 0.92 0.85 1.00 1.00 0.75 0.91
Graft loss, 2 years 0.94 0.97 1.00 0.94 0.90 1.00 0.96
Graft loss, 3 years 0.90 0.84 0.81 0.88 0.67 0.93 0.86
Graft loss, 5 years 0.98 0.91 0.81 1.00 1.00 0.94 0.95

ACC, accuracy; AUC, area under the curve; F1, F1-score; NPV, negative predictive value; PPV, positive predictive value; SENS, sensitivity; SPEC, specificity.

References

1. Webster AC, Nagler EV, Morton RL, Masson P. Chronic kidney disease. Lancet 2017;389:1238–1252.
crossref pmid
2. Romagnani P, Agarwal R, Chan JCN, et al. Chronic kidney disease. Nat Rev Dis Primers 2025;11:8.
crossref pmid pdf
3. Abecassis M, Bartlett ST, Collins AJ, et al. Kidney transplantation as primary therapy for end-stage renal disease: a National Kidney Foundation/Kidney Disease Outcomes Quality Initiative (NKF/KDOQITM) conference. Clin J Am Soc Nephrol 2008;3:471–480.
crossref pmid pmc
4. Hariharan S, Israni AK, Danovitch G. Long-term survival after kidney transplantation. N Engl J Med 2021;385:729–743.
crossref pmid
5. Majoni SW, Ullah S, Collett J, Hughes JT, McDonald S. Weight change trajectories in Aboriginal and Torres Strait islander Australians after kidney transplantation: a cohort analysis using the Australia and New Zealand Dialysis and Transplant registry (ANZDATA). BMC Nephrol 2019;20:232.
crossref pmid pmc pdf
6. Voora S, Adey DB. Management of kidney transplant recipients by general nephrologists: core curriculum 2019. Am J Kidney Dis 2019;73:866–879.
crossref pmid
7. Solez K, Axelsen RA, Benediktsson H, et al. International standardization of criteria for the histologic diagnosis of renal allograft rejection: the Banff working classification of kidney transplant pathology. Kidney Int 1993;44:411–422.
crossref pmid
8. Angelini A, Andersen CB, Bartoloni G, et al. A web-based pilot study of inter-pathologist reproducibility using the ISHLT 2004 working formulation for biopsy diagnosis of cardiac allograft rejection: the European experience. J Heart Lung Transplant 2011;30:1214–1220.
crossref pmid
9. Cheplygina V, de Bruijne M, Pluim JP. Not-so-supervised: a survey of semi-supervised, multi-instance, and transfer learning in medical image analysis. Med Image Anal 2019;54:280–296.
crossref pmid
10. Campanella G, Hanna MG, Geneslaw L, et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat Med 2019;25:1301–1309.
crossref pmid pmc pdf
11. Bilal M, Jewsbury R, Wang R, et al. An aggregation of aggregation methods in computational pathology. Med Image Anal 2023;88:102885.
crossref pmid
12. Ye Y, Xia L, Yang S, et al. Deep learning-enabled classification of kidney allograft rejection on whole slide histopathologic images. Front Immunol 2024;15:1438247.
crossref pmid pmc
13. Xiang H, Xiao Y, Li F, et al. Development and validation of an interpretable model integrating multimodal information for improving ovarian cancer diagnosis. Nat Commun 2024;15:2681.
crossref pmid pmc pdf
14. Naesens M, Roufosse C, Haas M, et al. The Banff 2022 Kidney Meeting Report: reappraisal of microvascular inflammation and the role of biopsy-based transcript diagnostics. Am J Transplant 2024;24:338–349.
crossref pmid
15. Yi Z, Xi C, Menon MC, et al. A large-scale retrospective study enabled deep-learning based pathological assessment of frozen procurement kidney biopsies to predict graft loss and guide organ utilization. Kidney Int 2024;105:281–292.
crossref pmid
16. Aziz F, Parajuli S, Jorgenson M, et al. Chronic active antibody-mediated rejection in kidney transplant recipients: treatment response rates and value of early surveillance biopsies. Transplant Direct 2022;8:e1360.
crossref pmid pmc
17. Gadermayr M, Tschuchnig M. Multiple instance learning for digital pathology: a review of the state-of-the-art, limitations & future potential. Comput Med Imaging Graph 2024;112:102337.
crossref pmid
18. Doupe P, Faghmous J, Basu S. Machine learning for health services researchers. Value Health 2019;22:808–815.
crossref pmid
19. Sageshima J, Than P, Goussous N, Mineyev N, Perez R. Prediction of high-risk donors for kidney discard and nonrecovery using structured donor characteristics and unstructured donor narratives. JAMA Surg 2024;159:60–68.
crossref pmid pmc
20. Focosi D, Vistoli F, Boggi U. Rejection of the kidney allograft. N Engl J Med 2011;364:485–486.
crossref
21. Kers J, Bülow RD, Klinkhammer BM, et al. Deep learning-based classification of kidney transplant pathology: a retrospective, multicentre, proof-of-concept study. Lancet Digit Health 2022;4:e18–e26.
crossref pmid
22. Wekerle T, Segev D, Lechler R, Oberbauer R. Strategies for long-term preservation of kidney graft function. Lancet 2017;389:2152–2162.
crossref pmid
23. Viglietti D, Loupy A, Aubert O, et al. Dynamic prognostic score to predict kidney allograft survival in patients with antibody-mediated rejection. J Am Soc Nephrol 2018;29:606–619.
crossref pmid


ABOUT
BROWSE ARTICLES
EDITORIAL POLICY
FOR CONTRIBUTORS
Editorial Office
#301, (Miseung Bldg.) 23, Apgujenog-ro 30-gil, Gangnam-gu, Seoul 06022, Korea
Tel: +82-2-3486-8736    Fax: +82-2-3486-8737    E-mail: registry@ksn.or.kr                

Copyright © 2026 by The Korean Society of Nephrology.

Developed in M2PI

Close layer