Introduction
Chronic kidney disease affects approximately 10% of the global population, with millions progressing to end-stage renal disease (ESRD) [
1,
2]. Kidney transplantation remains the preferred treatment for ESRD [
3]. However, posttransplant intrinsic kidney injury poses significant threats to graft survival, with rejection being the primary concern [
4]. Notably, kidney allograft rejection accounts for 63% of graft loss cases occurring >1 year posttransplantation [
5]. Therefore, an early and accurate diagnosis of rejection may optimize posttransplant management and prolong graft survival.
Histopathological assessment of allograft biopsies remains the cornerstone modality for diagnosing rejection [
6]. This process involves detecting and subtyping rejection as T cell-mediated rejection (TCMR), antibody-mediated rejection (ABMR), or mixed TCMR/ABMR (MIXR) [
7]. The prognosis of rejection is further stratified according to lesion severity. The subtyping and prognosis of rejection guide therapeutic decision-making and determine clinical management. However, manual assessment performed by transplant pathologists is inherently subjective, with inter-observer variability limiting reproducibility [
8]. Moreover, this process is time-consuming, expensive, and labor-intensive. Artificial intelligence (AI)-assisted digital pathology systems may overcome these limitations by providing more objective, efficient, and standardized analyses, and improving accuracy, consistency, and workflow.
Multiple instance learning (MIL) is a weakly supervised machine-learning paradigm that is well-suited for digital histopathological image analysis [
9,
10]. MIL effectively extracts discriminative features from whole-slide images (WSIs) using only slide-level labels, maintaining high diagnostic accuracy while significantly reducing the annotation burden [
11]. Recently, we developed MIL models based on WSIs for automated assessment of transplant pathology, specifically for kidney allograft rejection [
12]. However, these models can be further improved. First, previously proposed MIL models did not incorporate readily accessible clinicopathological variables that could aid in the diagnosis of rejection, including serum biomarkers such as creatinine (CR). Second, clinical decision-making is inherently a multimodal integration process [
13]. However, to the best of our knowledge, no previous study has systematically integrated multimodal information, including pathological images and clinicopathological texts, into a risk stratification framework for renal allograft rejection.
As such, the present study aimed to develop and validate “RenalTransPredNet,” an integrative multimodal information system, for the diagnosis of kidney allograft rejection. RenalTransPredNet combines WSIs with routine clinicopathological parameters to enhance predictive accuracy for the detection, subtyping, and prognosis of kidney allograft rejection. In addition to classifying lesions and predicting rejection prognosis, RenalTransPredNet provides a visual interpretation of the impact of each parameter on the predictions using Shapley additive explanation (SHAP) plots.
Discussion
Kidney transplant rejection exhibits highly complex features due to the intricate immune mechanism involved [
20]. Diagnosing rejection based solely on histopathological images remains challenging because differential diagnosis requires the integration of multimodal information. The diagnostic utility of image-based MIL predictions and clinical parameters has previously been explored for the diagnosis of kidney allograft rejection. However, the integration of multimodal features for improved diagnosis remains poorly understood. In this study, we introduced RenalTransPredNet, which integrates image-based MIL prediction and clinicopathological parameters to predict renal allograft rejection. For classifying kidney allograft lesions into four categories, RenalTransPredNet achieved higher AUC values than either single-modality approach, with an AUC of 0.886 in the internal testing set. Moreover, the model demonstrated predictive performance for treatment response in patients with rejection, as well as for assessing post-rejection graft loss across multiple time points. These results support the potential of RenalTransPredNet as an AI-assisted pathology analysis tool that enhances pathologists' accuracy in distinguishing between renal rejection and other allograft lesions, as well as in assessing rejection prognosis.
The strength of this study lies in the integration of image-based MIL prediction with routine clinicopathological parameter prediction to ensure complementary information and comprehensive feature representation. By classifying renal allograft lesions into four categories (i.e., TCMR/ABMR/MIXR/Other), RenalTransPredNet enabled the detection and subtyping of rejection and achieved a higher AUC value than the pathological image and clinicopathological information models in the independent internal testing set (AUC, 0.886 versus 0.831 or 0.801). Kers et al. [
21] constructed deep-learning models using WSIs to detect kidney rejection and reported that their model achieved an AUC of 0.78 in a multicenter setting. Ye et al. [
12] developed a histological image-based MIL model to classify renal rejection subtypes and achieved an AUC of 0.798 in a single-center setting. However, these studies relied solely on pathological images, which limited the comprehensive characterization of rejection features. We constructed RenalTransPredNet by integrating rejection characteristics across different dimensions, resulting in improved performance. These results confirm the feasibility of multimodal fusion for enhancing the model performance.
Additionally, we evaluated RenalTransPredNet for prognostic prediction in patients with rejection. In the internal testing sets, RenalTransPredNet achieved a high performance in terms of treatment response and graft loss. Both indicators were closely correlated with allograft prognosis. These indicators reflect kidney allograft function by assessing eGFR variations and dialysis resumption [
22,
23]. Here, we found that RenalTransPredNet demonstrated favorable performance for the 5-year survival of renal allografts after rejection (
Table 2). These results suggest that RenalTransPredNet may have the potential to contribute to prognostic assessments of rejection. Unfortunately, long-term follow-up data remain unavailable because the study was limited to cases enrolled over the previous decade.
Interpretability optimization was implemented using RenalTransPredNet. In our previous study, heat maps highlighted the key discriminative features of the pathological image model, including interstitial inflammation and tubulitis in TCMR, and peritubular capillaritis in ABMR [
12]. To identify key clinicopathological parameters, our statistical analysis of variables across different rejection subtypes and other lesions revealed significant differences in the following: age, CR variation ratio, months from transplantation to biopsy, C4d, DSA, and LYMPH. Although months from transplantation to biopsy differed significantly, correlation analysis revealed a significant association between months from transplantation to biopsy and chronicity scores (e.g., interstitial fibrosis [ci] and tubular atrophy [ct]) derived from WSIs (
Supplementary Fig. 3, available online). Thus, to avoid multicollinearity and improve the biological plausibility of predictions, our fusion model prioritized WSIs’ features that directly reflect tissue injury, while excluding the variable of months from transplantation to biopsy. The SHAP plots further show the relative contribution of each clinicopathological parameter to the RenalTransPredNet predictions. We observed that the CR variation ratio had the greatest effect on the decisions made by RenalTransPredNet for all clinicopathological parameters. This parameter reflects alterations in renal allograft function and serves as an indicator of rejection. Interestingly, it also plays a significant role in predicting the treatment response and renal allograft survival. C4d and LYMPH also made significant contributions in different predictive tasks. DSA and patient age had less of an impact on the decision. C4d and DSA are key elements in the pathological diagnosis of ABMR. However, DSA positivity was detected in 5.2% of patients (19/362) who underwent kidney allograft biopsies, which may partially account for the relatively low contribution of DSA. LYMPH reflects the patient’s immune status to some extent, which may influence their response to treatment and graft survival.
To evaluate whether our model provides independent diagnostic information beyond routine clinical markers, we stratified patients by renal function stability. Notably, even in patients with stable CR (variation <20%), the WSI-only model maintained a high discriminative ability (AUC, 0.817) and successfully identified subclinical rejection cases that would have been missed by CR monitoring alone. This suggests that WSI-derived morphological features can detect early rejection prior to significant functional deterioration, offering a window for timely intervention. Although further validation in larger cohorts is needed, these results provide evidence that our model adds independent and incremental value to conventional monitoring, potentially addressing an unmet need in non-invasive surveillance of transplant recipients.
The present study had some limitations. First, although RenalTransPredNet demonstrated high diagnostic performance across diverse clinical tasks, its generalizability is limited due to the lack of external validation. Additionally, as RenalTransPredNet was developed primarily using a cohort of clinical indication biopsies, its applicability to general screening or protocol biopsies requires further validation. Second, even in RenalTransPredNet, some cases with rejection were misclassified into other categories. As shown in the results, these misclassified cases shared common features: subtle morphological changes and atypical clinical presentations. This indicates that while RenalTransPredNet is highly sensitive to typical rejection, its sensitivity to minimal or subclinical lesions remains a limitation. Future work should include external multicenter validation and enlarge the training set with mild/borderline rejection cases to improve discrimination of subtle rejection phenotypes, including early ABMR. Third, other parameters, including donor characteristics, human leukocyte antigen mismatch counts, and immunosuppressive regimens, are important for clinical diagnostic tasks. Regrettably, retrospective studies have difficulty in collecting detailed information. The utilities of these parameters will be explored further in future prospective studies. Fourth, imaging examinations, such as ultrasound and magnetic resonance imaging, also play important roles in the diagnosis of kidney rejection, and the integration of these multimodal images may further improve the performance of RenalTransPredNet.
This study demonstrated the utility of RenalTransPredNet for the detection, subtyping, and prognostic prediction of kidney allograft rejection. RenalTransPredNet integrated pathological images with clinicopathological information and showed higher performance with an internal dataset, supporting the importance of multimodal integration in the diagnosis of kidney allograft rejection. The SHAP analysis further presented the interpretation of the decision of RenalTransPredNet and demonstrated the importance of each parameter. These findings suggest that RenalTransPredNet may have the potential to assist in improving pathologist detection and subtyping accuracy for renal allograft rejection and to contribute to prognostic assessments. Future prospective studies are needed to validate its clinical applicability.