Non-invasive predictive model for incidental gallbladder carcinoma based on multimodal features: Integrating clinical data, MRI radiomics, and deep transfer learning features.

OA: gold

Abstract

Gallbladder carcinoma, among the most prevalent malignancies of the biliary system, often presents with insidious early symptoms. Delayed diagnosis of incidental gallbladder carcinoma frequently leads to therapeutic delays, significantly compromising patient prognosis. Clinical data,preoperative Magnetic Resonance Imaging examinations, and histopathological records from incidental gallbladder carcinoma patients were retrospectively enrolled. An integrated noninvasive predictive model was developed by combining deep learning features extracted from Magnetic Resonance Imaging radiomics with key clinical predictors. Data from 299 patients with benign gallbladder disease and 106 incidental gallbladder carcinoma cases were analyzed. Multimodal deep learning algorithms identified an optimal predictive model. Multivariate analysis revealed hemoglobin, direct bilirubin, age, distance of the cystic duct from the confluence of right and left hepatic ducts, and diameter of the common bile duct as independent predictors of incidental gallbladder carcinoma (all P<0.01). The multimodal combined feature-based prediction model (AUC=0.894[95%CI 0.829-0.960]) surpassed unimodal models (Clinic:0.862 [ 95%CI 0.787-0.936]; Rad: 0.797 [ 95%CI 0.716-0.878]; DLR: 0.844[ 95%CI 0.775-0.912]) in test cohort. Haemoglobin, direct bilirubin,age, distance of the cystic duct from the confluence of right and left hepatic ducts, and diameter of the common bile duct constitute independent predictors of incidental gallbladder carcinoma. The multimodal combined feature-based prediction model significantly outperforms single-modality models, offering a robust tool for preoperative risk stratification.
Full text 57,145 characters · extracted from pmc-nxml · 10 sections · click to expand

Credit

Qiang Gao: Writing – review & editing, Writing – original draft, Conceptualization. Tiexin Liu: Methodology. Haoyu Song: Supervision. Yang Hu: Methodology. Shaohua Ren: Investigation. Zhenghao Li: Supervision. Huifang Yang: Resources. Liye Liu: Visualization. Xinyue Chang: Methodology. Yang Liu: Data curation. Fei Wang: Investigation. Defang Zhao: Supervision, Conceptualization. Xitao Liu: Conceptualization. Zhenxia Wang: Funding acquisition.

Ethics

All the study participants signed an informed consent for the inclusion in the study. The study was approved by the Research Ethics Committee of the Affiliated Hospital of Inner Mongolia Medical University (No: S.2025208). The procedures followed were in accordance with the ethical standards of the responsible committee on human experimentation (institutional or regional) and with the Helsinki Declaration of 1975, as revised in 2000.

Funding

This research is supported by Science and Technology Program of the Joint Fund of Scientific Research for the Public Hospitals of Inner Mongolia Academy of Medical Sciences (Grant Number: 2024GLLH0369 ) and Natural Science Foundation of Inner Mongolia Autonomous Region (Grant Number: 2025QN08020 ).

Results

All data analyses in this study were implemented in Python 3.7.12. Statistical analyses were performed using Statsmodels version 0.13.2. Radiomic feature extraction was achieved employing PyRadiomics 3.0.1. Machine learning models (including support vector machines [SVM]) were constructed upon Scikit-learn 1.0.2. The deep learning framework was developed using PyTorch 1.11.0, with GPU acceleration optimized through CUDA 11.3.1 and cuDNN 8.2.1. In this part of the analysis, we performed a between-group comparison of preoperative baseline clinical features of IGBC and Non-IGBC cases within the training and test cohorts. The results showed that both datasets were in age, weight, PMN (neutrophil count), LY (lymphocyte ratio), HGB (haemoglobin measurements), ALB (albumin), AST (albumin transaminase), TBIL (total bilirubin), DBIL, CBD diameter (common bile duct diameter), GBWET (gallbladder wall oedema thickness), GBV (gallbladder volume), GS (gallbladder stones) and CBDS (common bile duct stones) metrics were statistically significantly different (p<0.05). Of note, the training cohort additionally showed statistically significant (P<0.05) between-group differences in WBC (white blood cell count), BG (blood glucose) and CD-HDC distance (distance of the cystic duct from the confluence of the right and left hepatic ducts) parameters. Intergroup comparisons for the remaining baseline clinical parameters did not reach the threshold of statistical significance ( Table 1 ). One-way logistic regression analysis was performed on all preoperative clinical data, and variables with P1 and P<0.05. CBD diameter (OR:1.271, 95%CI [1.097-1.474], P=0.007), CD-HDC distance (OR: 1.082, 95%CI [1.038-1.129], P=0.002), AGE (OR:1.076, 95%CI [1.039-1.116], P=0.001) were established as independent risk factors for incidental gallbladder carcinoma ( Table 2 ). Table 2 Univariable and multivariable logistic regression analyses of clinical characteristics. Table 2 dummy alt text feature_name OR_UNI OR lower 95%CI_UNI OR upper 95%CI_UNI p_value_UNI OR_MULTI OR lower 95%CI_MULTI OR upper 95%CI_MULTI p_value_MULTI GS 0.549 0.447 0.675 0 0.724 0.391 1.342 0.389 CBDS 0.623 0.553 0.703 0 0.874 0.345 2.217 0.812 BG 0.875 0.839 0.912 0 1.344 1.025 1.761 0.072 PMN 0.884 0.843 0.927 0 1.533 0.779 3.019 0.299 WBC 0.897 0.867 0.928 0 0.688 0.378 1.25 0.303 CBD diameter 0.925 0.898 0.953 0 1.271 1.097 1.474 0.007 GS size 0.941 0.925 0.957 0 0.966 0.909 1.025 0.341 GBWET 0.954 0.914 0.996 0.07 LY 0.962 0.954 0.97 0 0.981 0.926 1.04 0.585 ALB 0.976 0.97 0.982 0 0.896 0.811 0.99 0.069 CD-HDC distance 0.98 0.972 0.988 0 1.082 1.038 1.129 0.002 LPS 0.98 0.973 0.986 0 0.998 0.985 1.011 0.762 CD-CBD distance 0.982 0.977 0.987 0 1.02 0.987 1.054 0.315 Weight 0.985 0.982 0.989 0 1 0.964 1.037 0.981 CD-CHD angle 0.987 0.983 0.991 0 1.006 0.995 1.016 0.377 LHD-RHD angle 0.989 0.986 0.992 0 1.001 0.987 1.014 0.944 Age 0.989 0.985 0.993 0 1.076 1.039 1.116 0.001 IBIL 0.989 0.978 0.999 0.084 AMY 0.991 0.987 0.994 0 0.997 0.983 1.011 0.741 HGB 0.993 0.991 0.995 0 0.963 0.933 0.994 0.049 LHD-CHD angle 0.993 0.992 0.995 0 0.993 0.977 1.008 0.423 RHD-CHD angle 0.994 0.992 0.995 0 1.011 0.994 1.028 0.278 PLT 0.996 0.995 0.997 0 0.998 0.992 1.003 0.496 AST 0.997 0.993 1 0.117 ALT 0.997 0.995 0.999 0.029 0.998 0.989 1.007 0.709 UAMY 0.997 0.996 0.998 0 0.998 0.997 1 0.115 GBV 1 1 1 0.805 TBIL 1.002 0.999 1.006 0.232 DBIL 1.007 1.002 1.013 0.032 1.05 1.009 1.092 0.042 Note:WBC:white blood cell count;PLT:Platelet Count;PMN:PolyMorphonuclear Neutrophil;LY:Lymphocyte ratio;HGB:Hemoglobin determination;ALB:Albumin;BG:Bloodglucose;ALT:Glutamic-pyruvic transaminase;AST:Aspartate aminotransferas;TBIL:Total bilirubin;DBIL:Direct Bilirubin;IBIL:Indirect Bilirubin;AMY:Amylase;UAMY:urinary amylase;LPS:Lipase;CBD diameter:common bile duct diameter;CD-HDC distance:Distance from the cystic duct to the confluence of the left and right hepatic ducts;CD-CBD distance:Distance from the cystic duct to the distal end of the common bile duct;GBWET:Edematous gallbladder wall thickness;CD-CHD angle:Angle of confluence between the cystic duct and common hepatic duct;LHD-RHD angle:Angle of confluence between the left and right hepatic ducts;LHD-CHD angle:Angle between the left hepatic duct and the common hepatic duct;RHD-CHD angle:Angle between the right hepatic duct and the common hepatic duct;GBV:Gallbladder volume;GS (Gallstone status classification):0: Solitary stone;1: Multiple stones;2: No stones (including cases with polyps or adenomyomatosis);3: Sludge/microlithiasis ;CBDS (Common Bile Duct Stone) classification:0: Solitary gallstone;1: Multiple gallstones;2: Stone-free status (including polyps or adenomyomatosis);3: Biliary sludge/microlithiasis Univariable and multivariable logistic regression analyses of clinical characteristics. A total of 311 features are extracted, and the number of each feature ①firstorder features: 198, glcm features:99, ③hape features:14. Radiomics features and deep learning features screening: after t-test, correlation coefficient and Least Absolute Shrinkage with Selection Operator (LASSO) regression analyses, significant radiomics features derived from radiomics and deep learning algorithms were screened ( Fig. 2 ). Fig. 2 10-fold cross-validated coefficients and mean squared error (MSE). A.The coefficient of 10 fold validation of imaging omics features B.Mean Squared Error (MSE) of RadiomicsFeatures C.Coefficients for 10 fold validation of imagingomics features combined with deep learning features D.MeanSquared Error (MSE) ofIntegrated Radiomics and Deep Learning Features. Fig 2 dummy alt text 10-fold cross-validated coefficients and mean squared error (MSE). A.The coefficient of 10 fold validation of imaging omics features B.Mean Squared Error (MSE) of RadiomicsFeatures C.Coefficients for 10 fold validation of imagingomics features combined with deep learning features D.MeanSquared Error (MSE) ofIntegrated Radiomics and Deep Learning Features. The final filtered radiomics feature parameters and corresponding feature weights as well as feature parameters and corresponding feature weights based on multimodal combined features ( Figure S1, Figure S2, and Table S2, Table S3.). Three types of machine learning discriminative models, namely Support Vector Machine (SVM), K-Nearest Neighbor (KNN) and Light Gradient Boosting Machine (LightGBM), were constructed based on the feature screening, and their predictive performance are systematically evaluated and compared. The study was evaluated by multidimensional performance indicators, including: diagnostic accuracy, area under the working characteristic curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), precision, recall and F1-Score were quantitatively analysed in the training cohort and independent test cohort. The results show that all three types of algorithmic models exhibit good discriminative performance. In view of the limited sample size, the training cohort with the advantage of data size was used as the benchmark,and LightGBM was finally established as the optimal algorithm. The predictive efficacy of LightGBM for Incidental Gallbladder Carcinoma in the training cohort with the three feature sets was as follows: (i) AUC=0.862 (95%) CI 0.787-0.936); (ii) AUC=0.951 (95%CI 0.923-0.980); and (iii) AUC=0.986(95%CI 0.976-0.996)( Fig. 3 , Table 3 , Table 4 , Table 5 ). LightGBM was established as the optimal algorithm in this study. Its selection was based not only on achieving the highest AUC value in the training cohort but also on its superior comprehensive performance: (1) High computational efficiency, supporting rapid training and hyperparameter tuning on large-scale feature sets; (2) As a gradient-boosting algorithm based on tree models, it automatically handles complex nonlinear relationships among features, aligning closely with our objective of multimodal feature fusion to capture intricate interactions; (3) As shown in Fig. 3 and Table 3 , Table 4 , Table 5 , the LightGBM demonstrated relatively more robust generalization performance across most feature sets in the test cohort, exhibiting smaller performance degradation compared to SVM and KNN; (4) As demonstrated by the decision curve analysis (DCA) in Fig. 6 , the LightGBM consistently outperforms other models in terms of clinical net benefit, directly validating its clinical decision-making value. Furthermore, the calibration curve of the LightGBM ( Fig. 6 ) exhibits excellent calibration, indicating high consistency between predicted and true probabilities, which further supports its status as a reliable predictive tool. Fig. 3 Area under the curve (AUC) for the predictive performance of each model in the training and test sets. A.ROC Curve of the Machine Learning Model for Clinical Features in the training set B.ROC Curve of theMachine Learning Model for Clinical Features in the test set C,ROC Curve of the Radiomics Feature-Based MachineLearning Model in the training set D.ROC Curve of the Radiomics Feature-Based Machine Learning Model in the testset E.ROC Curve of the Machine Learing Model for Early Fusion (Radiomics +Deep Learning) Features in the trainingset F .ROC Curve of the Machine Learing Model for Early Fusion (Radiomics + Deep Learning) Features in the test set. Fig 3 dummy alt text Table 3 Performance comparison of machine learning models for clinical feature-based prediction. Table 3 dummy alt text Accuracy AUC 95% CI Sensitivity Specificity PPV NPV Precision Recall F1 Threshold Task SVM 0.782 0.889 0.8446 - 0.9327 0.857 0.751 0.583 0.929 0.583 0.857 0.694 0.196 label-train SVM 0.84 0.884 0.8180 - 0.9492 0.778 0.857 0.609 0.931 0.609 0.778 0.683 0.287 label-test KNN 0.835 0.916 0.8834 - 0.9482 0.843 0.832 0.67 0.929 0.67 0.843 0.747 0.4 label-train KNN 0.87 0.844 0.7666 - 0.9224 0.556 0.96 0.8 0.883 0.8 0.556 0.656 0.6 label-test LightGBM 0.897 0.949 0.9244 - 0.9744 0.9 0.896 0.778 0.957 0.778 0.9 0.834 0.333 label-train LightGBM 0.833 0.862 0.7875 - 0.9358 0.75 0.857 0.6 0.923 0.6 0.75 0.667 0.315 label-test Table 4 Comparison of machine learning predictive models based on radiomics features. Table 4 dummy alt text Accuracy AUC 95% CI Sensitivity Specificity PPV NPV Precision Recall F1 Threshold Task SVM 0.922 0.962 0.9341 - 0.9896 0.929 0.919 0.823 0.97 0.823 0.929 0.872 0.294 label-train SVM 0.704 0.798 0.7142 - 0.8814 0.833 0.667 0.417 0.933 0.417 0.833 0.556 0.238 label-test KNN 0.831 0.922 0.8902 - 0.9545 0.871 0.815 0.656 0.94 0.656 0.871 0.748 0.4 label-train KNN 0.704 0.732 0.6381 - 0.8262 0.667 0.714 0.4 0.882 0.4 0.667 0.5 0.4 label-test LightGBM 0.905 0.951 0.9232 - 0.9796 0.886 0.913 0.805 0.952 0.805 0.886 0.844 0.319 label-train LightGBM 0.71 0.797 0.7161 - 0.8779 0.806 0.683 0.42 0.925 0.42 0.806 0.552 0.29 label-test Table 5 Comparison of machine learning predictive models based on early fusion features. Table 5 dummy alt text Accuracy AUC 95% CI Sensitivity Specificity PPV NPV Precision Recall F1 Threshold Cohort SVM 0.996 0.998 0.9953 - 1.0000 0.986 1 1 0.994 1 0.986 0.993 0.369 label-train SVM 0.722 0.809 0.7244 - 0.8933 0.861 0.683 0.437 0.945 0.437 0.861 0.579 0.121 label-test KNN 0.893 0.943 0.9164 - 0.9689 0.829 0.919 0.806 0.93 0.806 0.829 0.817 0.4 label-train KNN 0.778 0.742 0.6420 - 0.8424 0.611 0.825 0.5 0.881 0.5 0.611 0.55 0.4 label-test LightGBM 0.93 0.986 0.9756 - 0.9959 0.943 0.925 0.835 0.976 0.835 0.943 0.886 0.312 label-train LightGBM 0.759 0.844 0.7753 - 0.9118 0.806 0.746 0.475 0.931 0.475 0.806 0.598 0.219 label-test Area under the curve (AUC) for the predictive performance of each model in the training and test sets. A.ROC Curve of the Machine Learning Model for Clinical Features in the training set B.ROC Curve of theMachine Learning Model for Clinical Features in the test set C,ROC Curve of the Radiomics Feature-Based MachineLearning Model in the training set D.ROC Curve of the Radiomics Feature-Based Machine Learning Model in the testset E.ROC Curve of the Machine Learing Model for Early Fusion (Radiomics +Deep Learning) Features in the trainingset F .ROC Curve of the Machine Learing Model for Early Fusion (Radiomics + Deep Learning) Features in the test set. Performance comparison of machine learning models for clinical feature-based prediction. Comparison of machine learning predictive models based on radiomics features. Comparison of machine learning predictive models based on early fusion features. To explore the recognition ability of the deep learning model for different samples, this study used the Gradient-weighted Class Activation Mapping (Grad-CAM) technique for visualisation and analysis. As shown in the figure below, Grad-CAM reveals the key image regions on which the model decision depends by highlighting the significant activation regions of the final convolutional layer associated with cancer type prediction, thus providing a basis for the interpretability of the model ( Fig. 4 ). Fig. 4 Class activation mapping(CAM). Fig 4 dummy alt text Class activation mapping(CAM). After the construction and validation of the machine learning prediction model, each feature set showed good fitting performance in both the training and test cohorts. Specifically, in the training cohort, the clinical feature parameter set, the radiomics feature parameter set, and the deep learning-derived radiomics (DLR) features parameter set show excellent fitting performance; and in the test cohort, similar generalisation performance to the training cohort is observed. In the study,a logistic regression algorithm was used to implement a feature pre-fusion strategy for the clinical features and the deep learning-derived radiomics (DLR) features, construct a linear prediction model (Nomogram, see Fig. 5 below), and systematically evaluate the performance differences of each feature set in the LightGBM classifier. The results show that the multimodal combined features parameters set achieves the optimal discriminative performance in the training cohort, with an area under the curve (AUC) of 0.995; in the test cohort, the area under the curve (AUC) of 0.894. To further assess the incremental value of the combined model, pairwise DeLong tests, IDI, and NRI analyses were performed. In the test cohort, the combined model achieved the highest AUC among all models. DeLong analysis showed that the combined model significantly outperformed the radiomics-only model, while IDI analysis demonstrated improved discrimination compared with the clinical, radiomics, and DLR models. In addition, all NRI values were positive, indicating improved reclassification performance of the combined model (Figure S3, Figure S4, Figure S5). The calibration curve analysis shows a significant consistency between the model's predicted probability and the measured probability, suggesting that the model is well calibrated, and the Hosmer–Lemeshow goodness-of-fit test showed acceptable calibration of the Combined model in both the training cohort (p = 0.057) and the test cohort (p = 0.204) (Table S4). The model was validated by decision curve analysis (DCA), which showed that the model maintained a high net benefit over a wide range of threshold probabilities, confirming its outstanding clinical applicability ( Table 6 ). Fig. 5 Nomogram. Note: HGB;Hemoglobin determination; DBlL:Direct Biliru bin; CD-HDC distance: Distance from the cystic duct tothe confluence of the left and right hepatic ducts, CBD diameter: common bile duct diameter;, DLR:Deep Learning-Integ!ated Radiomics. Fig 5 dummy alt text Table 6 Comparison of linear models based on early fusion features. Table 6 dummy alt text Signature Accuracy AUC 95% CI Sensitivity Specificity PPV NPV Cohort Clinic 0.897 0.949 0.9244 - 0.9744 0.9 0.896 0.778 0.957 train Rad 0.905 0.951 0.9232 - 0.9796 0.886 0.913 0.805 0.952 train DLR 0.93 0.986 0.9756 - 0.9959 0.943 0.925 0.835 0.976 train Combined 0.959 0.995 0.9903 - 0.9998 0.986 0.948 0.885 0.994 train Clinic 0.833 0.862 0.7875 - 0.9358 0.75 0.857 0.6 0.923 test Rad 0.71 0.797 0.7161 - 0.8779 0.806 0.683 0.42 0.925 test DLR 0.759 0.844 0.7753 - 0.9118 0.806 0.746 0.475 0.93 test Combined 0.889 0.894 0.8291 - 0.9597 0.722 0.937 0.765 0.922 test Note: Clinic:Clinica features; Rad:Radiomic features; DLR:Deep Learning-Integrated Radiomics; Combined:Combined Feature Fusion at Input Level: Integrating Clinical, Radiomic, and Deep Learning-derived Features. Nomogram. Note: HGB;Hemoglobin determination; DBlL:Direct Biliru bin; CD-HDC distance: Distance from the cystic duct tothe confluence of the left and right hepatic ducts, CBD diameter: common bile duct diameter;, DLR:Deep Learning-Integ!ated Radiomics. Comparison of linear models based on early fusion features. Calibration curves are a core tool for assessing the quality of model prediction probabilities in deep learning, and the calibration curves in this study visually demonstrate that the confidence level of the model predictions has a very high level of real confidence. Decision curve analysis (DCA) results show that the multimodal combined features model has a significant advantage in terms of prediction probability and consistently exhibits higher net return potential under different risk thresholds compared to other feature models, further confirming the effectiveness of the model ( Fig. 6 ). Fig. 6 Model evaluation in training and test sets: AUC, calibration curve, and decision curve analysis. A.ROC-AUC Performance on the Independent training set B.ROC-AUC Performance on the lndependent test seiC.Calibration curve of the training group D.Calibration curve of the test group E.Decision curve analysis of the traininggroup F.Decision curve analysis of the test group Note: Clinic:Clinica features:Rad:Radiomic features, DLR:Deep Learning-Integrated Radiomics; Combined:CombinedFeature Fusion at Input Level: Integrating Clinical, Radiomic, and Deep Learning-derived Features. Fig 6 dummy alt text Model evaluation in training and test sets: AUC, calibration curve, and decision curve analysis. A.ROC-AUC Performance on the Independent training set B.ROC-AUC Performance on the lndependent test seiC.Calibration curve of the training group D.Calibration curve of the test group E.Decision curve analysis of the traininggroup F.Decision curve analysis of the test group Note: Clinic:Clinica features:Rad:Radiomic features, DLR:Deep Learning-Integrated Radiomics; Combined:CombinedFeature Fusion at Input Level: Integrating Clinical, Radiomic, and Deep Learning-derived Features.

Technical

This study proposes a non-invasive predictive model based on clinical-omics features, radiomics features, and deep learning features for predicting incidental gallbladder cancer. First, manual features are extracted from Magnetic Resonance Cholangiopancreatography (MRI) images using the PyRadiomics library. Next, depth features are obtained by cropping the maximum slice of regions of interest (ROIs) in MRI images, enhanced through transfer learning using a pre-trained ResNet18 model to improve feature expression. Through correlation filtering and Lasso regression, robust, low-redundancy features with predictive value were selected. Ultimately, a multidimensional feature model and corresponding nomogram were established in an independent validation cohort. The figure below illustrates the radiomics analysis workflow adopted in this study ( Fig. 1 ). Fig. 1 Study flow diagram. Fig 1 dummy alt text Study flow diagram. (1) Patients who underwent cholecystectomy from January 2014 to December 2024 at the First Affiliated Hospital of Inner Mongolia Medical University; (2) Patients with no preoperative diagnosis of gallbladder malignancy; (3) Postoperative pathology confirmed a malignant tumour of the gallbladder. (4) Age 15-94. (1) Patients with a clear preoperative diagnosis of a malignant tumour of the gallbladder; (2) There was no complete preoperative MRI. (3) Lack of preoperative laboratory tests (blood count, liver function, coagulation function, etc.). Through the electronic medical record system of the Affiliated Hospital of Inner Mongolia Medical University, we collected patients ' basic information, general laboratory indicators, surgery-related indicators, tumour-related indicators, and MRI imaging data. The basic information included: Gender, age, weight, and postoperative recovery time. General laboratory indicators: including blood routine, liver function, MRI imaging-related indicators: we collected preoperative MRI imaging data of the patients (The sixth slice of the patient's MRCP images was measured. If this specific slice was not clearly discernible, the most distinct slice from the MRCP series was selected and analyzed instead.), including common bile duct diameter (CBD diameter), distance of the cystic duct from the confluence of the right and left hepatic ducts(CD-HDC distance), distance of the cystic duct from the end of the common bile duct (CD-CBD distance), thickness of gallbladder wall oedema (GBWET), gallbladder volume (GBV), angle of confluence of the cystic duct and common hepatic duct (CD-CHD angle), angle between the left and right hepatic ducts (LHD-RHD angle), angle between the left hepatic duct and the common hepatic duct (LHD-CHD angle), angle between the right hepatic duct and the common hepatic duct (RHD-CHD angle), and the presence or absence of stones in the gallbladder (GS, 0: single stone1: multiple stones 2: no stones or polyps, adenomyosis 3: sedimentary stones), whether there are stones in the common bile duct (CBDS, 0: single stone 1: multiple stones 2: no stones or polyps, adenomyosis 3: sedimentary stones), the size of gallbladder stones (GS size), and the size of the common bile duct stones (CBDS size), (averaged over two measurements) were measured by 3D-slicer ( www.slicer.org ) software using Total Segmentation plug-in to segment and calculate the gallbladder volume, and the features of interest of the gallbladder and biliary system were segmented separately by ITK-SNAP, and the values of the features of interest were measured separately. A total of 405 patients were included in this part of the study based on the initial screening population, inclusion criteria, and exclusion criteria. Randomisation was done in the ratio of 7:3, 284 in the training cohort, of which 74 were patients with incidental gallbladder carcinoma, and 121 in the test cohort, of which 31 were patients with incidental gallbladder carcinoma. MRI Data Acquisition and Sequence Selection: A total of 405 MRI scans were retrospectively included in this study. T2-weighted imaging (T2WI) was selected as the sole input sequence, given its clinical significance and superior soft tissue contrast in identifying pathological lesions [ 16 , 17 ]. All images were acquired following standard clinical imaging protocols on a clinical MRI scanner, with consistent acquisition parameters (e.g., slice thickness, field of view) across all subjects. ROI sketching: In this study,the region of interest (ROI) was manually sketched based on the training dataset by ITK-SNAPv3.8.0software for gallbladder ROIs, which were labelled independently by two experienced senior physicians, and in case of disagreement, a senior physician with more than 20 years of experience was used for arbitration. This delicate and strict manual sketching process provides a high-quality training database for subsequent segmentation model development. To ensure methodological rigour and reproducibility in radiomics feature extraction, we quantified operator-dependent variability using the intraclass correlation coefficient (ICC). Two senior abdominal radiologists (with 8 and 10 years of experience, respectively) who participated in ROI segmentation independently annotated 30 randomly selected MRI cases: Intra-observer agreement: The primary radiologist repeated segmentation on the same cases after a 2-week interval. Inter-observer agreement: Both radiologists independently segmented the identical case set. A two-way random-effects absolute agreement model (ICC(2,1)) was implemented via their package in R. Results demonstrated: Median intra-observer ICC: 0.92 (range: 0.81–0.98); Median inter-observer ICC: 0.88 (range: 0.75–0.96). Geometric features (ICC=0.94), intensity features (ICC=0.95), and texture features (ICC=0.93) exhibited optimal stability. All operations were performed on standardised MRI sequences (T2-weighted imaging) following a unified protocol: Systematic exclusion of pericholecystic non-target tissues (e.g., liver parenchyma, bowel); Consistent boundary definition using ITK-SNAP's preset 5-mm ROI expansion function. This workflow adhered to Radiomics Quality Score (RQS v2.0) guidelines, significantly reducing operator-dependent variability. Features with ICC >0.75 were retained for subsequent analysis (detailed categorical data: Supplementary Table S1). Feature extraction: manual features are classified into three categories:(I) geometric features, (II) intensity features, and (III) texture features. Geometric features describe the three-dimensional morphological properties of the tumour,reflecting its spatial structure;intensity features portray the homogeneity or heterogeneity within the tumour through the first-order statistical properties of the voxel grey scale; texture features are extracted from the voxel grey scale covariance matrix (GLCM), autocorrelation, cluster prominence, cluster shade, and cluster tendency, contrast, correlation, joint average, joint energy, joint entropy and other second and higher order statistical methods to quantify the spatial and spatial characteristics of grey intensity. methods to quantify the spatial distribution pattern of grey intensity and reveal the complex spatial correlation between voxels. A total of 311 manual features were extracted, including 14 geometric features, 99 intensity features and 198 texture features. All manual features wereextracted by an in-house feature analysis program developed based on the PyRadiomics library ( http://pyradiomics.readthedocs.io ).( Table S3 ) Feature screening: in this study, all extracted features were standardised by Z-score and assessed for statistical significance by t-test, and only features with p-value less than 0.05 were retained. To reduce covariance, the correlation between features was assessed using Pearson's correlation coefficient, and if the absolute value of the correlation coefficient between two features was≥0.9, one of the features was excluded. The optimal regularisation parameter was further determined by LASSO regression combined with 10-fold cross-validation to filter out the most predictive and informative subset of features. ROI cropping: in the methodological process of this study, for each patient's MRI image, T2-weighted imaging (T2WI) sequences were used as the input modality for deep learning feature extraction, consistent with the ROI segmentation protocol described in Section 2.2.1. A single slice at the largest level of the region of interest (ROI) was selected as a representative image. To simplify the complexity of the algorithm and reduce background interference, only the smallest outer rectangular region containing the ROI was retained. Considering that 2D features are generally more robust, while 3D approaches are more prone to overfitting in small-sample (n=74 IGBC cases in training cohort) settings and 2.5D methods provide limited additional benefit [ 18 , 19 ], a 2D-based feature extraction strategy was adopted in this study. Data enhancement: in this study, the RGB channel intensity distributions of the images are normalised by Z-score normalisation and the normalised images are used as model inputs. In the training phase, a real-time data enhancement strategy is used, and for the test images, only normalisation is performed. Data Normalization: We normalized the grayscale values of the image slices using a min-max transformation to adjust the range to [-1. 1]. Each croppedsubregion image was resized to 224×224 pixels using nearest neighbor interpolation, ensuring compatiblity with the input requirements of our deep learning models. Transfer learning: The ResNet18 architecture, pre-trained on the large-scale ImageNet dataset, was adopted as the feature extraction backbone. The model was trained in an end-to-end manner for 100 epochs on our target dataset, allowing all convolutional and fully connected layers to be adaptively updated. After training, high-level discriminative features were extracted from the final average pooling (avgpool) layer of the ResNet18 network, producing a 512-dimensional feature vector for each MRI image. These features were subsequently used as the input for the subsequent prediction model. This study explores the effectiveness of ResNet18 mainstream network architectures in optimising the performance of traditional CNN models, and conducts a comparative analysis of these models in order to filter the algorithms that best fit the needs of this study. Hyperparameters: to ensure that the model maintains its validity across different patient populations with significant heterogeneity, this study employs a transfer learning strategy to initialise the model using pre-training weights from the ImageNet database in order to enhance its adaptability to diverse datasets. The key lies in fine-tuning the learning rate to promote the generalisation performance of the model across datasets, for which we adopt the cosine decay learning rate strategy, whose formula is defined as follows: ηt=ηmin+12 (ηmax-ηmin)(1+cos (tTmaxπ)) ηt=ηmin+21 (ηmax-ηmin)(1+cos (Tmaxtπ)) Where ηmax=0.001, ηmax=0.001 is the initial learning rate, ηmin=0.0001, ηmin=0.0001 is the minimum learning rate, t is the number of current training rounds (epoch), Tmax=100, Tmax=100 is the total number of training rounds. Translation of the formula: ηt=ηmin+21 (ηmax-ηmin)(1+cos (Tmaxtπ)) Hyperparameter description: the symbolηmin=0, ηmin=0 represents the minimum learning rate andηmax=0.01, ηmax=0.01 represents the maximum learning rate. The parameter Ti=30, Ti=30 represents the number of cycles (epoch) in the iterative training process. Other key hyperparameters include the use of stochastic gradient descent (SGD) as the optimiser and the use of softmax cross entropy as the loss function. Radiomics features: by using the LASSO algorithm for strict feature screening, this study constructed a radiomics risk model using machine learning algorithms such as logistic regression, support vector machine and random forest. This ensemble approach enhances model robustness by aggregating predictions across diverse training configurations. The performance of each model is evaluated through comparative analysis, and the benefits of multimodal feature fusion are further explored by integrating multimodal features, thus exploring the advantages of multimodal combined model in improving prediction accuracy. Deep Learning-Derived Radiomics (DLR) features: In this study, for the optimal deep learning model (the CNN model performed best in the test cohort), we utilised its penultimate layer for feature extraction. Given the complexity of the CNN model, we reduced these features to 30 dimensions by principal component analysis (PCA). The DLR feature construction followed a systematic fusion pipeline: (1) Feature concatenation: deep learning features (30 dimensions post-PCA) were concatenated with radiomics features (311 dimensions) to form a unified feature pool of 341 dimensions; (2) Dimensionality reduction and redundancy removal: the concatenated features underwent correlation filtering (Pearson's correlation coefficient threshold |r|≥0.9 for redundancy elimination), followed by LASSO regression with 10-fold cross-validation to select the most predictive feature subset; (3) Feature re-ranking: the retained features were re-ranked based on their LASSO coefficients to establish the final feature hierarchy; (4) Unified model training: the refined DLR feature set was used to train integrated prediction models (SVM, KNN, LightGBM) on the combined training cohort, yielding the DLR score subsequently incorporated into the multimodal nomogram. Multimodal combined features: in order to enhance the clinical applicability of the model, we performed univariate and stepwise multivariate analyses on all clinical features to screen important variables. A logistic regression (LR) linear model was constructed by integrating the screened clinical features with the predictions of the deep learning model, resulting in a multimodal combined features, which was visually presented using a column-line diagram. To assess the adequacy of the logistic regression models relative to sample size, the events per variable (EPV) ratio was calculated based on the 74 IGBC events in the training cohort. For the multivariable clinical logistic regression model (5 independent predictors: HGB, DBIL, Age, CD-HDC distance, and CBD diameter), EPV = 74/5 = 14.8. For the multimodal combined logistic regression model (6 inputs: the above 5 clinical predictors plus the DLR score), EPV = 74/6 ≈ 12.3. Both values exceed the commonly recommended minimum threshold of 10, indicating that both logistic regression models are statistically supported by adequate sample sizes [ 20 , 21 ]. To avoid data leakage, all feature screening, dimensionality-reduction, and model-development procedures were performed exclusively within the training cohort. Specifically, t-test filtering, Pearson correlation-based redundancy removal, LASSO feature selection, PCA dimensionality reduction, model fitting, and hyperparameter tuning were conducted using only the training cohort. The independent test cohort was not involved in these procedures and was used only for final model evaluation. In this study, the discriminative efficacy of the deep learning model was rigorously assessed in the test cohort by plotting the subject's work characteristics (ROC) curves, analysing the calibration of the model using calibration curves and the Hosmer–Lemeshow goodness-of-fit test. In addition, the clinical utility of the predictive model was assessed by Decision Curve Analysis (DCA), thus clarifying the value of the model's potential benefit in clinical practice. The Shapiro-Wilk test was used in this study to verify the normality of the clinical features. Continuous variables were analysed by choosing the t-test or Mann-Whitney U-test according to their distributional characteristics, and categorical variables were analysed using the chi-square (χ²) test. The baseline features of all cohorts are detailed in Table 1 . Notably, the p-values for inter-cohort comparisons were all greater than 0.05, suggesting that the differences between cohort were not statistically significant, thus confirming the unbiased nature of the groupings. Table 1 Baseline clinical characteristics of benign gallbladder diseases and incidental gallbladder carcinoma in the training and validation cohorts. Table 1 dummy alt text feature_name Training Cohort: Non-Incidental GBC Training Cohort: Incidental GBC pvalue Validation Cohort: Non-Incidental GBC Validation Cohort:Incidental GBC pvalue AGE 53.32±13.58 67.90±10.85 <0.001 53.50±12.45 68.39±11.12 <0.001 Weight 67.88±11.59 61.48±10.88 <0.001 67.26±13.08 61.18±9.69 0.018 WBC 6.47±2.77 7.09±2.96 0.034 6.21±2.31 6.80±2.72 0.21 PLT 246.82±68.67 240.16±65.77 0.671 248.05±63.89 212.61±74.85 0.008 PMN 4.13±2.61 5.11±2.87 <0.001 3.80±1.99 4.57±2.53 0.034 LY 30.28±10.09 22.45±10.88 <0.001 31.64±9.28 25.53±10.87 <0.001 HGB 147.20±17.32 133.41±19.63 <0.001 145.22±14.67 129.86±22.91 <0.001 ALB 43.67±4.28 38.63±5.32 <0.001 43.00±3.31 38.25±5.65 <0.001 BG 5.24±1.36 5.95±2.22 0.003 5.27±1.75 5.79±2.00 0.219 ALT 50.93±87.71 71.56±95.97 0.184 46.86±85.20 113.12±146.29 0.005 AST 30.40±47.66 52.27±59.72 <0.001 27.12±41.34 67.00±67.01 <0.001 TBIL 19.12±16.95 66.50±107.16 0.043 16.56±10.81 92.69±121.62 <0.001 DBIL 7.79±9.40 46.36±82.24 0.032 7.02±6.76 70.93±103.39 <0.001 IBIL 11.34±10.70 19.23±26.88 0.274 11.40±24.66 21.45±20.77 0.009 AMY 91.24±235.78 73.40±94.86 0.062 63.72±21.12 60.25±16.34 0.791 UAMY 367.50±528.95 277.90±216.63 0.618 278.24±238.22 256.75±103.55 0.845 LPS 75.60±305.98 36.25±17.29 0.21 38.93±20.13 35.59±11.61 0.931 CBD diameter 6.52±2.20 8.53±3.93 <0.001 6.54±2.19 8.81±2.98 <0.001 CD-HDC distance 24.98±9.74 33.01±11.89 <0.001 24.76±9.17 27.78±12.45 0.221 CD-CBD distance 48.55±12.93 47.51±15.48 0.59 48.48±13.75 54.78±13.77 0.016 GBWET 3.45±1.77 6.48±3.14 <0.001 3.75±1.51 5.98±3.00 <0.001 CD-CHD angle 61.02±35.69 54.54±40.81 0.051 64.49±37.71 63.16±46.91 0.539 LHD-RHD angle 77.97±30.12 74.07±39.84 0.343 167.06±926.11 66.69±41.28 0.009 LHD-CHD angle 137.66±22.39 132.04±30.95 0.5 137.59±21.08 137.72±30.85 0.502 RHD-CHD angle 134.13±23.76 136.85±24.61 0.301 132.40±26.16 131.20±31.76 0.807 GBV 52565.48±39943.68 123629.39±139467.19 <0.001 51065.38±31659.76 103188.39±68236.90 <0.001 GS size 13.69±7.43 12.49±6.32 0.288 13.04±7.33 15.74±9.20 0.137 GS <0.001 <0.001 0 39(22.54) 13(18.57) 25(19.84) 5(13.89) 1 118(68.21) 46(65.71) 88(69.84) 22(61.11) 2 null 9(12.86) null 9(25.00) 3 16(9.25) 2(2.86) 13(10.32) null CBDS <0.001 <0.001 0 4(2.31) null null 3(8.33) 1 1(0.58) 9(12.86) null 6(16.67) 2 167(96.53) 60(85.71) 126(100.00) 27(75.00) 3 1(0.58) 1(1.43) null null Note:WBC:white blood cell count;PLT:Platelet Count;PMN:PolyMorphonuclear Neutrophil;LY:Lymphocyte ratio;HGB:Hemoglobin determination;ALB:Albumin;BG:Bloodglucose;ALT:Glutamic-pyruvic transaminase;AST:Aspartate aminotransferas;TBIL:Total bilirubin;DBIL:Direct Bilirubin;IBIL:Indirect Bilirubin;AMY:Amylase;UAMY:urinary amylase;LPS:Lipase;CBD diameter:common bile duct diameter;CD-HDC distance:Distance from the cystic duct to the confluence of the left and right hepatic ducts;CD-CBD distance:Distance from the cystic duct to the distal end of the common bile duct;GBWET:Edematous gallbladder wall thickness;CD-CHD angle:Angle of confluence between the cystic duct and common hepatic duct;LHD-RHD angle:Angle of confluence between the left and right hepatic ducts;LHD-CHD angle:Angle between the left hepatic duct and the common hepatic duct;RHD-CHD angle:Angle between the right hepatic duct and the common hepatic duct;GBV:Gallbladder volume;GS size:Gallstone size;GS (Gallstone status classification):0: Solitary stone;1: Multiple stones;2: No stones (including cases with polyps or adenomyomatosis);3: Sludge/microlithiasis ;CBDS (Common Bile Duct Stone) classification:0: Solitary gallstone;1: Multiple gallstones;2: Stone-free status (including polyps or adenomyomatosis);3: Biliary sludge/microlithiasis Baseline clinical characteristics of benign gallbladder diseases and incidental gallbladder carcinoma in the training and validation cohorts. Data Analysis Tools: All data analyses in this study were implemented based on Python 3.7.12 environment. Statistical analyses were done using Statsmodels 0.13.2, and radiomics feature extraction was implemented through PyRadiomics 3.0.1. Machine learning models (e.g. Support Vector Machine SVM,etc.) were constructed based on Scikit-learn 1.0.2, and the deep learning framework was developed relying on PyTorch 1.11.0 and accelerated and optimised by CUDA 11.3.1 and cuDNN 8.2.1. For multiple performance metrics involved in machine learning model comparisons, the focus of reporting lies on point estimates (e.g. AUC) and their confidence intervals. We mitigated overfitting risks through cross-validation and independent test set evaluations, without performing multiple comparison corrections between models. Pairwise DeLong tests, IDI, and NRI analyses were additionally performed to evaluate the incremental predictive value of the combined model.

Conclusion

This study successfully achieved its objectives by constructing a high-performance model for preoperative non-invasive prediction of incident gallbladder cancer (IGBC). Through systematic analysis, we first identified key clinical risk factors including hemoglobin, direct bilirubin, age, istance of the cystic duct from the confluence of right and left hepatic ducts (CD-HDC distance), and diameter of the common bile duct (CBD). Subsequently, we developed prediction models based on clinical features, MRI radiomics features, and deep learning features, respectively. Ultimately, the multimodal combined model integrating all dimensions proved to be the optimal solution, achieving exceptionally high predictive performance (AUC=0.894) on the test set and significantly outperforming any single-feature model. Thus, this study not only validates the superiority of multimodal data fusion strategies but also provides a powerful tool for the clinical implementation of precise preoperative identification of UGBC.

Discussion

In recent years, studies have shown that radiomics has demonstrated significant clinical application potential in early screening,pathological staging, and prognostic assessment of gallbladder carcinoma by deeply mining the texture,morphology, and functional heterogeneity information implicit in multimodal imaging data, such as CT and MRI, and by using non-invasive and reproducible high-throughput quantitative analysis techniques. It has been reported in the literature that based on the preoperative enhanced CT data of 168 gallbladder carcinoma patients from 2011 to 2020,the study integrated the radiomics feature (1502→13) and clinical variables through nnU-Net segmentation and DeepSurv model to construct a survival prediction model, whose C-index and 1/2/3-year AUC were comparable to the manual segmentation model (Delong test P>0.05), which can be preoperatively assessed for individualised survival stratification and guide diagnostic and treatment decisions [ 22 ]. Another study found that based on the radiomics and clinical data of 195 patients with gallbladder carcinoma in two tertiary hospitals in Shanghai, the study constructed a multimodal survival prediction model by extracting features through 3D-DenseNet and combining Cox regression, which had a C-index of 0.787, and the AUC of 1/3/5-year survival rates were 0.827, 0.865, and 0.926, which were significantly better than the single-peak model and the traditional TNM staging system, which can improve the accuracy of prognostic assessment [ 23 ]. Deep learning refers to a set of computer models that have achieved breakthroughs in recent years in the way computers extract information from images. These algorithms have been applied to a number of medical speciality tasks, including radiomics and pathology, and their performance has been comparable to the level of senior physicians in specific scenarios. A Transformer architecture that combines global and local attention mechanisms and bag-of-words feature embedding has been proposed in the literature for high-precision detection of gallbladder carcinoma (GBC) from ultrasound images, with an accuracy that exceeds that of radiologists and is highly interpretable, validating diagnostic features consistent with the medical consensus as well as discovering new visual markers, which can be used as a second reading tool to aid diagnosis [ 24 ]. A novel deep neural network combining global and local attention mechanisms and Transformer architecture has been reported to detect gallbladder carcinoma (GBC) with high accuracy from ultrasound images, which exceeds the accuracy of human radiologists, and interpretable analysis based on bag-of-words embedding verifies the consistency of the model's decision-making with the medical consensus and reveals potential new diagnostic features, which can be used as a secondary reading tool for assisted diagnosis [ 25 ]. Given the advantages of deep learning in medicine, our algorithm for deep learning was used in this study, and several predictive models constructed had significant predictive efficacy. Previous studies have shown that arsenic compounds in blood can be molecularly docked with haemoglobin, and their elevated concentrations are positively correlated with incidence gallbladder carcinoma [ 26 ]. This finding suggests that haemoglobin status may be potentially associated with gallbladder carcinoma. Analysis of our central cohort showed that lower haemoglobin levels were significantly associated with increased incidence of incident gallbladder carcinoma (IGBC). We hypothesise that this negative association stems from the profound impact of IGBC on the nutritional status of patients compared to benign gallbladder disease: patients with IGBC often present with more severe anorexia and progressive weight loss, leading to increased anaemia, which further reduces haemoglobin concentration. Meanwhile, the prediction model constructed in our centre showed that elevated direct bilirubin was significantly associated with the risk of incident gallbladder carcinoma (IGBC). Recent studies have further confirmed that direct bilirubin is a key biomarker (p<0.05) for differentiating yellow granulomatous cholecystitis (XGC) from gallbladder carcinoma [ 27 ]. The underlying mechanism is that progression of IGBC can lead to biliary obstruction, contributing to significantly higher direct bilirubin levels than in patients with benign gallbladder disease. This feature provides an important basis for clinical differentiation between IGBC and benign lesions. Aging exhibits a significant age-dependent effect in incidental gallbladder carcinoma (IGBC). Sharma et al. demonstrated that the development of gallbladder carcinoma is closely associated with the accumulation of age-associated mutations, and that aging is a central driver of cancer evolution [ 28 ]. Our central cohort further revealed that ageing significantly increases the risk of IGBC, and the underlying pathological basis may lie in the fact that chronic mechanical irritation triggered by benign lesions such as long-term gallbladder stones accelerates the malignant transformation of the gallbladder epithelium through the recurrent mucosal damage-repair cycle. Bray's global epidemiological study similarly confirmed that the incidence of gallbladder carcinoma increased exponentially among people≥65 years of age [ 29 ], which is highly in line with the present study findings are highly consistent. Matsubara et al.demonstrated that abnormal biliopancreatic duct confluence is a high-risk factor for gallbladder carcinoma, suggesting that anatomical variations of the biliary tract may drive the cancerous process [ 30 ]. In the present study,we innovatively proposed the choledochal duct-common hepatic duct confluence distance (CD-HDC distance) as a novel predictive marker and found that prolongation of this distance was significantly and positively associated with the risk of incidental gallbladder carcinoma (IGBC). The underlying mechanism may involve:prolongation of the CD-HDC distance→disturbance of biliary fluid dynamics→chronic nonstruvous cholecystitis→progressive biliary epithelial heteroplasia→ultimately driving malignant transformation. The biological rationale lies in the fact that an extended CD-HDC distance may alter local bile fluid dynamics. A longer confluence pathway may lead to impaired bile emptying, creating relative obstruction and bile stasis, thereby triggering chronic non-calculous cholecystitis. This prolonged chronic inflammatory stimulus, through a "damage-repair-proliferation" cycle, ultimately drives dysplasia and malignant transformation of the gallbladder mucosal epithelium. This hypothesis is pathophysiologically consistent with Matsubara et al.'s findings that abnormal hepatopancreatic duct confluence promotes gallbladder cancer development [ 30 ]. Regarding the reproducibility of this feature measurement, this study ensured reliability through a rigorous process: two senior pathologists, unaware of pathological outcomes, independently performed measurements on MRCP sequences using ITK-SNAP software. This approach has been widely validated in clinical imaging studies to enhance measurement accuracy and consistency. Although MRI scanning protocols may vary across institutions, key biliary structures typically appear stable on MRCP sequences, providing a foundation for standardized CD-HDC distance measurements across centers. Future multicenter studies will focus on evaluating the reproducibility of this feature between observers and scanning devices. Takuma et al. found, based on a nationwide study in Japan, that the incidence of biliary tract cancer in patients with pancreaticobiliary meridian abnormality (PBM) with bile duct dilatation reached 10.6%, of which gallbladder carcinoma accounted for 64.9%, while the proportion of gallbladder carcinoma in those with pancreaticobiliary meridian abnormality (PBM) without bile duct dilatation was as high as 93.2% [ 31 ]. Our centre further revealed that reduced CBD diameter was significantly negatively associated with the risk of incidental gallbladder carcinoma (IGBC) by standardised measurement of common bile duct diameter (CBD diameter). This finding is anatomically corroborated with the"high prevalence of gallbladder carcinoma in non-dilated PBM"in Takuma's study [ 31 ], in which a reduction in CBD diameter may reflect relative obstruction of the cystic ducts, leading to the accumulation of bile in the gallbladder and accelerating the process of localised cancerous transformation. The multimodal combined model developed in this study demonstrated outstanding performance in the training cohort (AUC=0.995). Although its performance decreased in the test cohort (AUC=0.894), its discriminatory capability remains significantly superior to conventional clinical assessment, indicating clear potential for clinical translation. We envision this model being integrated into radiology information systems as a preoperative decision-support tool in real-world clinical settings. When patients undergo cholecystectomy for suspected benign gallbladder disease, their clinical data and preoperative MRI images can be input into the model to generate an individualized probability of IGBC. This probability value can be combined with existing standards such as imaging reports and tumor markers to provide surgeons with more comprehensive risk stratification. For example, in patients with high model-predicted probabilities, even if conventional imaging does not suggest malignancy, preparations for laparoscopic-to-open conversion during the initial surgery, extended gallbladder bed resection, or intraoperative frozen section examination could be recommended. This approach avoids secondary surgeries due to missed diagnoses, directly improving surgical radicality and patient prognosis. Decision curve analysis (DCA) in this study demonstrates that the combined model delivers higher clinical net benefit across a broad range of threshold probabilities. This decision science validation confirms its clinical utility, indicating that adopting this model to guide clinical decisions can yield greater benefits for patient populations. This study has the following limitations: First, as mentioned in section 2.4, the sample size of this study was sufficient for logistic regression modeling, but it was still limited by the extremely low incidence of incidental gallbladder cancer [ 32 , 33 ]. The gap between training AUC (0.995) and test AUC (0.894) reflects a degree of overfitting due to the limited training sample.Furthermore, the overlapping confidence intervals across models (combined: 95%CI 0.829–0.960; clinical: 95%CI 0.787–0.936; radiomics: 95%CI 0.716–0.878; DLR: 95%CI 0.775–0.912) suggest that performance differences may not reach formal statistical significance. Second, the absence of external validation prevents comprehensive assessment of the model's universality and robustness across different institutions and scanning protocols. Third, preoperative diagnostic heterogeneity may introduce labeling bias, as some cases with atypical malignant features on preoperative imaging were included, potentially confusing the features learned by the model. To address these limitations, future research plans include: First, initiating multicenter external validation as the primary task to overcome current constraints. Second, employing advanced algorithmic techniques to tackle the small-sample challenge, such as incorporating stricter regularization during model training, applying oversampling for class imbalance, and utilizing Bootstrap resampling to more robustly estimate model performance confidence intervals. Third, we will advance cross-modal deep integration by exploring the fusion of liver-specific contrast-enhanced MRI sequence features and liquid biopsy molecular biomarkers (e.g., ctDNA) to construct more biologically meaningful predictive models.

Introduction

Incidental gallbladder carcinoma (IGBC) refers to gallbladder cancer incidentally discovered during pathological examination after cholecystectomy performed for benign conditions, with an incidence ranging from 0.2% to 2.8% [ 1 ]. Due to its asymptomatic nature preoperatively and the lack of sensitive diagnostic tools [ 2 ], IGBC is frequently missed, leading to delayed treatment and significantly worsened prognosis [ 3 ]. With the widespread adoption of laparoscopic cholecystectomy, the detection rate of IGBC has been increasing [ 4 ], yet its postoperative recurrence risk remains high [ 5 ], making it a formidable clinical challenge that urgently requires the development of effective preoperative noninvasive prediction methods. Current preoperative precision diagnosis of IGBC faces substantial difficulties. Conventional imaging modalities (e.g., ultrasound, CT, MRI) exhibit limited discriminatory power for early-stage lesions and are prone to confusion with benign conditions [ 6 , 7 ]. Simultaneously, serum tumour markers (e.g., CA19-9) demonstrate low sensitivity in early stages [ 8 , 9 ]. Although emerging radiomics and deep learning technologies have demonstrated potential in tumour diagnosis by extracting high-throughput features from medical images beyond human visual recognition [ [10] , [11] , [12] , [13] , [14] , [15] ] significant knowledge gaps remain in systematically integrating these techniques with clinical characteristics to construct a comprehensive model dedicated to preoperative prediction of IGBC. To address this gap, this study aims to develop and validate a multimodal predictive model for achieving noninvasive preoperative risk assessment of IGBC. Specifically, we will integrate three dimensions of data: patient clinical characteristics, MRI-based radiomics features, and deep learning features. By constructing and comparing single-feature models with fusion models, this study seeks to determine the optimal predictive combination, ultimately providing clinicians with a robust decision support tool to improve early diagnosis and treatment strategies for IGBC patients.

Coi Statement

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data Availability

All data are available from the corresponding author with reasonable request. Supplementary materials. Table S1. Inter- and intra-observer reproducibility assessment of radiomic features using intraclass correlation coefficient (ICC). Table S2. Radiomic Feature Coefficients in Radiomics Model. Table S3. Multimodal feature coefficients in radiomics model. Table S4. Hosmer–Lemeshow goodness-of-fit test p-values for model calibration. Figure S1. Radiomic Features. Figure S2. multimodal combined features. Figure S3. IDI analysis in the training and test cohorts. Figure S4. NRI analysis in the training and test cohorts. Figure S5. Pairwise DeLong test results in the training and test cohorts.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-16T09:21:09.727480+00:00