Explainable Multimodal Deep Learning Model for Early Prediction of Treatment-Requiring Retinopathy of Prematurity

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background: Retinopathy of prematurity (ROP) is a leading preventable cause of childhood blindness. Current screening guidelines, based primarily on gestational age and birth weight, result in numerous unnecessary examinations. We aimed to develop an explainable multimodal deep learning model for early prediction of treatment-requiring ROP. ‏ Methods: In a retrospective cohort of 384 preterm infants (203 treated, 181 untreated) from a tertiary center in Iran (2021-2024), we integrated four directional fundus images, semi-supervised vessel segmentation maps (using a U-Net-based adversarial domain-adaptation approach), and comprehensive clinical/demographic data. A multi-view fusion model with an attention mechanism extracted vessel-aware features, reduced via PCA, and combined with clinical variables. Six multimodal feature sets were evaluated using 14 machine learning classifiers with 5-fold stratified cross-validation. ‏ Results: The best-performing models (Extra Trees and Random Forest) on the full multimodal feature set achieved a test accuracy of 0.987, an AUC-ROC of 0.999, an F1-score of 0.988, and near-perfect specificity (up to 1.000). Interpretability analyses (SHAP and Grad-CAM) confirmed that predictions were primarily driven by vascular morphology features (PCA1) and posterior pole abnormalities consistent with plus disease. ‏ Conclusions: The proposed explainable multimodal model significantly outperforms clinical-only approaches and represents a promising tool for risk stratification in ROP screening. It has the potential to reduce unnecessary examinations, infant stress, and healthcare burden while facilitating timely intervention. External multicenter validation is warranted.
Full text 138,551 characters · extracted from preprint-html · click to expand
Explainable Multimodal Deep Learning Model for Early Prediction of Treatment-Requiring Retinopathy of Prematurity | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Explainable Multimodal Deep Learning Model for Early Prediction of Treatment-Requiring Retinopathy of Prematurity Fatemeh Sefidbaf, Fatemeh Baharvand Ahmadi, Nasser Shoeibi, Saeid Eslami, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8838478/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 14 You are reading this latest preprint version Abstract Background: Retinopathy of prematurity (ROP) is a leading preventable cause of childhood blindness. Current screening guidelines, based primarily on gestational age and birth weight, result in numerous unnecessary examinations. We aimed to develop an explainable multimodal deep learning model for early prediction of treatment-requiring ROP. ‏ Methods: In a retrospective cohort of 384 preterm infants (203 treated, 181 untreated) from a tertiary center in Iran (2021-2024), we integrated four directional fundus images, semi-supervised vessel segmentation maps (using a U-Net-based adversarial domain-adaptation approach), and comprehensive clinical/demographic data. A multi-view fusion model with an attention mechanism extracted vessel-aware features, reduced via PCA, and combined with clinical variables. Six multimodal feature sets were evaluated using 14 machine learning classifiers with 5-fold stratified cross-validation. ‏ Results: The best-performing models (Extra Trees and Random Forest) on the full multimodal feature set achieved a test accuracy of 0.987, an AUC-ROC of 0.999, an F1-score of 0.988, and near-perfect specificity (up to 1.000). Interpretability analyses (SHAP and Grad-CAM) confirmed that predictions were primarily driven by vascular morphology features (PCA1) and posterior pole abnormalities consistent with plus disease. ‏ Conclusions: The proposed explainable multimodal model significantly outperforms clinical-only approaches and represents a promising tool for risk stratification in ROP screening. It has the potential to reduce unnecessary examinations, infant stress, and healthcare burden while facilitating timely intervention. External multicenter validation is warranted. Retinopathy of Prematurity multimodal analysis machine learning deep learning fundus images Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Retinopathy of Prematurity (ROP) is a disease characterized by abnormal retinal blood vessel growth in preterm newborns [ 1 ]. The condition typically starts when these infants are exposed to high oxygen levels (hyperoxia), causing existing vessels to close and regress. When oxygen levels normalize again, the resulting relative hypoxia drives excessive new vessel formation, mainly through vascular endothelial growth factor (VEGF) [ 2 ]. Known risk factors include extremely low gestational age or birth weight, inflammation during pregnancy or after birth, the general immaturity of premature infants, lung-related complications, and anemia [ 3 ]. Each year, around 15 million premature infants are born globally, placing approximately 1.4 million at risk for ROP [ 4 ]. Roughly 32,000 of these infants progress to severe visual impairment or blindness [ 4 ]. In Iran, prevalence rates reported in various studies range widely from 5.6% to 70.3% [ 5 – 30 ]. Available treatments include laser photocoagulation, cryotherapy, and anti-VEGF injections. Still, early diagnosis via timely screening is essential for successful management [ 31 , 32 ]. Current guidelines rely mostly on gestational age and birth weight, which frequently leads to screening many infants who will not develop treatment-requiring disease [ 33 – 35 ]. Frequent eye examinations can be stressful and uncomfortable for the infants, while also creating emotional and financial strain for families [ 36 , 37 ]. Factors like delayed referrals, missed appointments, shortage of pediatric retinal specialists, and inadequate parental awareness further raise the chances of preventable blindness [ 4 ]. Recent studies have shown accuracies over 90% for image-only models in detecting plus disease and cases that need intervention. These results come from applying convolutional neural networks to fundus photographs (Shoeibi et al., 2025[ 38 ]; Wang et al., 2021[ 39 ]). Multimodal approaches – the ones that mix retinal images with clinical and demographic data – usually do better than image-only models. This advantage is clearest in predicting disease progression or reactivation after anti-VEGF treatment (Wu et al., 2025[ 40 ]). Researchers have also come up with explainable AI methods. These provide clearer interpretations of ROP stages and key pathological features, which helps increase trust among clinicians and makes adoption easier (Wang et al., 2021[ 39 ]). On the other hand, systematic reviews make it clear that many of these models still have drawbacks: not enough external validation, datasets that are too small, and limited inclusion of wider clinical factors. All this points to a continuing need for stronger, more widely applicable, and truly interpretable multimodal systems (Jafarizadeh et al., 2025[ 41 ]). The present study introduces an explainable multimodal deep learning model. It combines retinal images with a broad range of clinical and demographic variables – things like infant sex, parental education, pregnancy and birth characteristics, neonatal factors (including NICU stay duration), and eye-specific severity measures – to predict treatment-requiring ROP. The model is designed to enhance risk stratification, allow earlier detection of high-risk cases, support timely treatment, minimize unnecessary examinations, reduce healthcare system burden, and ultimately improve care quality and blindness prevention in this vulnerable population. Methods 1. Study Design and Participant Cohort This retrospective cohort study was conducted at Khatam-al-Anbia Ophthalmology Hospital, the primary tertiary referral center for neonatal ophthalmic care in northeastern Iran. The study was approved by the Mashhad University of Medical Sciences Ethics Committee (IR.MUMS.MEDICAL.REC.1403.338) and adhered to the tenets of the Declaration of Helsinki. Infants were systematically referred for retinopathy of prematurity screening according to national guidelines, which specify eligibility as gestational age < 34 weeks and/or birth weight ≤ 2000 g. Consecutive infants presenting to the dedicated ROP clinic between January 2021 and December 2024 were included. All infants underwent standardized comprehensive ROP evaluation using the RetCam imaging system, documenting ROP stage, zone, and presence of plus disease. Infants with complete evaluation data were eligible; those with unresolved missing data on key variables were excluded. Treatment eligibility (laser photocoagulation or intravitreal anti-VEGF injection) was independently assessed by two senior pediatric ophthalmologists with > 10 years of experience in ROP management. Discrepancies were adjudicated by a third senior ophthalmologist, yielding a kappa coefficient of 0.89 (95% CI 0.85–0.93). From the screened population, 400 treatment-requiring infants were identified. A comparison group of 400 infants without treatment indication was randomly selected. Following data quality assessment, the final analytical cohort comprised 384 infants (203 treated, 181 untreated) with complete multimodal datasets. Written informed consent was waived by the ethics committee due to the retrospective design and use of fully anonymized data. 2. Data Acquisition and Preprocessing 2.1. Clinical and Demographic Variables Clinical and demographic data were extracted from electronic health records using Python 3.12 with pandas. Variables included infant sex, twin status, ROP features at first visit (stage, zone, and plus disease), gestational age and birth weight, age and weight at first ROP visit, NICU stay duration, parental ages at birth, parental education levels, gravidity, pregnancy initiation mode, pre-pregnancy maternal status, and city of residence. After data cleaning, no missing values remained in the analyzed variables. Numerical variables underwent outlier removal (interquartile range method) and standardization (using StandardScaler fitted on the training set). Categorical variables were label-encoded (using LabelEncoder fitted on the training set). The target variable was binary treatment requirement (1 = treatment required, 0 = no treatment). 2.2. ROP image selection and preprocessing 2.2.1. Image Quality Control All retinal images underwent a rigorous quality control (QC) procedure prior to inclusion in the analytical pipeline to ensure reliability of downstream segmentation, feature extraction, and classification tasks. The heterogeneous nature of neonatal fundus imaging—characterized by variable illumination, motion artifacts, media opacity, and differences in field-of-view—necessitates stringent QC to prevent poor-quality images from degrading model performance or introducing bias. Therefore, a standardized multi-stage QC framework was implemented to systematically evaluate the suitability of each image. First, all fundus photographs were screened by trained graders for essential diagnostic visibility, including adequate depiction of the posterior pole, optic disc, and major vascular arcades. Images with severe blur, defocus, or motion streaks that obscured vascular structures were excluded. Particular attention was given to vessel clarity and contrast, as vessel segmentation accuracy and subsequent vascular feature extraction depend heavily on the visibility of fine vascular structures. Images with excessive glare, overexposure, or shadowing caused by eyelids or specular reflections were removed if these artifacts interfered with vessel morphology assessment. Cases with incomplete field-of-view that did not capture the posterior pole or displayed only peripheral retina were also excluded, as these images do not provide sufficient information for evaluating plus disease or predicting treatment-requiring ROP. Following the manual grading step, automated quality metrics were applied to ensure consistency across the dataset. These included evaluations of illumination uniformity, contrast distribution, edge sharpness, and signal-to-noise ratio. Images failing predefined thresholds were flagged for secondary review and excluded when quality was deemed insufficient for reliable vessel segmentation. After completing all QC stages, the four highest-quality fundus images for each eye—captured from four standard directional views—were selected as the final candidates for subsequent processing phases, ensuring that downstream algorithms operated on the most informative and diagnostically reliable representations of the retina. This hybrid QC approach—combining expert review with automated assessment—ensured that only images meeting minimum diagnostic and computational quality requirements were used for model development. 2.2.2 Semi-supervised Vessel Segmentation Given the scarcity of expertly annotated retinal vessel masks and the abundance of unlabeled ROP fundus images typically encountered in clinical settings, we adopted a semi-supervised segmentation strategy inspired by the Deep Adversarial Network framework of Zhang et al. [ 42 ]. As illustrated in Fig. 1 , the model employs a U-Net architecture with a ResNet-34 encoder pretrained on ImageNet to enhance feature generalization and training stability. In the supervised warm-up phase, the network was trained using high-quality labeled retinal vessel datasets (DRIVE, STARE, CHASEDB1), with a combined Dice and binary cross-entropy loss. In the second stage, a dual-branch encoder–decoder model was introduced to jointly process both labeled target-domain images and unlabeled source-domain images. The public datasets were defined as the target domain, representing high-quality labeled vascular structures, whereas the locally acquired ROP images constituted the source domain, reflecting real-world clinical variability. During this phase, the segmentation branch of the network was frozen to preserve the vascular knowledge learned from the target domain, while the reconstruction branch remained trainable to encourage the extraction of domain-invariant retinal representations. Adversarial and reconstruction-based consistency constraints were applied to ensure that the feature distributions of source-domain images gradually aligned with those of the target domain. Importantly, the classifier associated with this stage was trained to output True for target-domain images and False for source-domain images, enabling the model to explicitly learn domain discrimination. This domain-classification signal plays a key role in guiding the encoder toward representations that minimize domain shift, thereby improving robustness in subsequent clinical prediction tasks. In the third stage of the framework, the dual-branch encoder–decoder architecture operates under an adversarial domain-alignment objective. At this point, the reconstruction branch of the network is frozen, preserving the domain-invariant structural representations learned in previous phases, while the segmentation branch remains fully trainable to allow targeted refinement of vessel probability maps for source-domain ROP images. Simultaneously, the domain-classification module of the framework is also frozen, ensuring that its decision boundary remains fixed during this phase. The segmentation branch is then optimized such that its outputs for source-domain images are increasingly likely to be interpreted by the frozen classifier as target-domain segmentation maps. This setup forces the segmentation network to adapt its output distribution toward that of the labeled target-domain datasets, thereby reducing residual domain discrepancies. As a result, the segmentation model learns to generate vessel maps for heterogeneous, real-world ROP images that closely resemble the structured vascular patterns seen in high-quality public datasets, achieving robust domain-invariant segmentation performance. 3. DL model development As illustrated in Fig. 2 , The development of the proposed deep learning framework proceeded in two main stages, beginning with a single-image classification model and subsequently extending to a multi-view architecture capable of leveraging four directional fundus images from each eye. The goal of this staged development was to establish a robust baseline classifier and then enhance it through the incorporation of semi-supervised segmentation priors and multi-task feature extraction. In the first stage, a standard image-based classifier was trained to predict treatment-requiring ROP from a single posterior-pole fundus image. A convolutional neural network pretrained on ImageNet served as the encoder backbone, enabling the model to benefit from strong visual feature representations despite limited domain-specific data. The classifier was optimized end-to-end using cross-entropy loss, and this single-image model provided an essential baseline for understanding the discriminative retinal features associated with clinically significant ROP. In the second stage, the framework was expanded to incorporate all four available fundus images captured from different viewpoints around the optic disc. For each view, vessel segmentation probability maps generated from the preceding semi-supervised segmentation model were used as additional structural inputs. Each fundus image and its corresponding vessel map were passed through one of four encoder streams, all sharing the same architecture but initialized with pretrained weights from the earlier training phase. To accommodate the single-channel vessel segmentation maps while preserving the pretrained structure of the encoder, the first convolutional layer was adapted by computing the channel-wise mean of its pretrained ImageNet filters and applying this averaged kernel to the segmentation input; after this modification of the initial layer, all subsequent convolutional blocks remained identical to the pretrained backbone. Thus, the fundus images were processed using the original three-channel pretrained weights, whereas the segmentation masks were integrated through a minimally adjusted first layer followed by the same deeper encoder architecture. 4. Multimodal Fusion and Machine Learning Evaluation The multimodal fusion framework integrates four directional fundus images and their corresponding vessel segmentation maps with structured demographic and clinical information to generate comprehensive representations for predicting treatment-requiring ROP. Each fundus image is paired with its vessel probability map and processed jointly through one of four parallel encoder streams with shared architecture. These encoders are initialized with weights pretrained in earlier model development stages, allowing extraction of stable, vessel-aware feature embeddings across retinal viewpoints. Independent processing of each directional view captures spatially distributed vascular cues—including tortuosity, dilation, branching complexity, and disc-centered geometry—that vary with camera angle and illumination conditions. The four feature vectors are projected into a shared latent space and fused using a multi-head attention module that adaptively weights the contribution of each view, prioritizing informative perspectives while suppressing noisier inputs. This produces a robust fused imaging representation. PCA dimensionality reduction is applied to this representation to generate low-dimensional embeddings (PCA1–3). All clinical and demographic variables included in the study, as well as the composition of the multimodal feature sets, were selected based on clinical expertise from experienced pediatric ophthalmologists. Selection prioritized variables with established or potential associations with ROP severity and treatment requirement. The fused imaging representation (including PCA1–3) is concatenated with clinical/demographic variables to form five predefined multimodal feature sets for comparative evaluation: Set 1: PCA1–3, ROP stage, zone, plus disease at first visit. Set 2: Clinical/demographic variables (birth age/weight, age/weight at first visit, NICU duration, parental age/education, sex, twin status, pregnancy initiation, pre-pregnancy issues, gravidity, city of residence). Set 3: Set 2 + PCA1–3. Set 4: Set 2 + ROP stage, zone, plus disease at first visit. Set 5: Set 2 + PCA1–3 + ROP stage, zone, plus disease at first visit. Set 6: All features (Set 1 + Set 2) Data were split 80/20 (stratified by target, random_state = 42). Preprocessing was fitted on the training set only. Given the mild class imbalance in the dataset (Class Ratio 0:1 = 0.89:1, corresponding to approximately 89 non-treatment cases per treatment case), class weighting was applied during model training to mitigate bias toward the majority class without altering the original data distribution. Specifically, the class_weight=balanced parameter was used in tree-based ensembles (Random Forest, Extra Trees, XGBoost, Gradient Boosting, AdaBoost, Decision Tree) and equivalent weighting strategies were implemented in other classifiers (e.g., class_weight=balanced in Logistic Regression and SVM). This approach was preferred over resampling techniques (e.g., oversampling or undersampling) due to its computational efficiency, preservation of the original data distribution, and avoidance of potential artifacts from synthetic sample generation in a clinically sensitive and moderately sized dataset. The following machine learning models were evaluated on each feature set: Random Forest (n_estimators = 200, max_depth = 10, class_weight='balanced'), XGBoost (n_estimators = 150, max_depth = 6, learning_rate = 0.1), Gradient Boosting, Logistic Regression (C = 0.1, penalty='l2'), SVM (kernel='rbf', C = 1.0, gamma='scale'), MLP (hidden_layers=(50,25), alpha = 0.0001), AdaBoost, Extra Trees, Decision Tree, KNN (n_neighbors = 5), Gaussian Naive Bayes, LDA, and QDA. By combining multi-view vascular morphology with clinical context across different feature sets, the framework enables comparative assessment of multimodal contributions to treatment prediction. 5. Statistical Analysis of Feature Importance Univariate statistical analyses were performed to assess the importance and association of individual features with treatment requirement. Descriptive statistics (mean, SD, median, min/max, skewness, kurtosis for numerical; frequency/percentage for categorical) were computed. Normality was tested using Shapiro-Wilk/Kolmogorov-Smirnov. Group comparisons used automated test selection: Student's t-test/Welch's t-test for normal data, Mann-Whitney U for non-normal. Categorical associations used χ² or Fisher's exact test with Cramér's V. Pearson correlation was calculated for numerical variables with treatment. These statistical tests provided evidence of feature importance without formal automated selection. Results 1. Semi-supervised Vessel Segmentation Figures 3 and 4 present qualitative comparisons of retinal vessel segmentation outputs generated by the proposed semi-supervised framework across multiple ROP cases, each containing four directional fundus views. For each sample, the original ROP images are shown in the top row, followed by the vessel maps produced using an unsupervised baseline method and the refined vessel masks obtained through the proposed semi-supervised segmentation model. These visual results demonstrate the substantial improvements achieved when unlabeled ROP images are leveraged through adversarial and reconstruction-based consistency training. Across all samples, the unsupervised segmentation outputs exhibit common failure patterns, including fragmented vessel structures, loss of fine-caliber peripheral vessels, and incomplete delineation near the optic disc region. These limitations are particularly evident in fundus views with uneven illumination, motion blur, or glare—conditions frequently encountered in neonatal imaging. Additionally, unsupervised masks often show excessive noise or spurious edges in low-contrast regions, reducing their suitability for downstream vascular feature extraction. By contrast, the semi-supervised segmentation maps display markedly enhanced anatomical coherence and vessel continuity across all four views. The proposed method successfully recovers subtle vascular branches, improves delineation of tortuous vessels, and reduces background noise, even in images with challenging acquisition conditions. The integration of labeled public fundus datasets during supervised warm-up, followed by domain adaptation on unlabeled ROP images, enables the model to generalize effectively to neonatal retinal morphology, which differs substantially from adult fundus patterns. This is particularly notable in peripheral zones where vessel caliber is small and prone to segmentation errors in conventional unsupervised methods.Moreover, the semi-supervised approach yields consistent vessel structures across all views within each sample, demonstrating the model’s robustness to variations in angle, illumination, and retinal curvature. This cross-view consistency is essential because the segmentation outputs are subsequently used as structural priors in the multi-view classification framework. Accurate preservation of vascular geometry, especially dilation, branching density, and tortuosity is critical for downstream prediction of treatment-requiring ROP. The improved vessel fidelity achieved via semi-supervised segmentation substantially enhances the reliability of extracted vascular biomarkers. 2. Model Interpretability Analysis To evaluate the clinical validity and transparency of the proposed predictive framework, we performed an extensive interpretability analysis using gradient-based class activation mapping (Grad-CAM) applied to the final multimodal fusion model. The goal of this analysis was to determine whether the model relies on physiologically meaningful retinal features—particularly vascular abnormalities, when predicting treatment-requiring ROP. Figure 5 presents representative examples of the original fundus images, their corresponding vessel segmentation maps, and the model-generated attention heatmaps. Across all samples, the visual explanations consistently demonstrated a strong alignment between the model’s decision-making process and established clinical hallmarks of severe ROP. In nearly all cases, the Grad-CAM activation maps concentrated on the posterior pole, including the optic disc and major vascular arcades. This corresponds precisely to the anatomical region where plus disease—the strongest predictor of treatment requirement—is assessed by clinicians. Increased tortuosity and dilation of central vessels are central diagnostic indicators in the ICROP classification system, and the model’s selective attention to these regions suggests that it has successfully learned vascular patterns associated with advanced disease. Notably, the heatmaps rarely emphasized peripheral retinal regions, which aligns with clinical practice, as treatment decisions are largely driven by posterior vascular abnormalities rather than the peripheral stage alone. This focal attention on the posterior pole provides strong evidence that the model’s predictions are grounded in pathophysiologic cues rather than spurious correlations. Overlaying the heatmaps onto the segmentation maps further confirmed that the model’s attention aligns with specific vascular structures rather than non-informative image regions such as illumination artifacts, imaging borders, or peripheral blur. In multiple examples, activation clusters followed the trajectory of dilated or tortuous vessels, indicating that the model leverages vessel morphology—rather than global image color or brightness—to guide its predictions. In cases with highly engorged vessels or asymmetric vascular branching, the heatmaps highlighted exactly those aberrant areas, demonstrating sensitivity to clinically meaningful nuances. For images with milder vascular changes, the attention maps were more diffuse yet still centered around the arcades, reflecting the model’s uncertainty in borderline cases—behavior that mimics real-world variability among human graders. Importantly, the interpretability results reveal consistency between the vessel-aware representations learned during the semi-supervised segmentation stage and the downstream classification behavior. Because the fusion model integrates both raw fundus features and segmentation-derived vascular context, the resulting attention maps capture a blend of structural and morphological cues. This demonstrates that the segmentation-guided feature extraction process not only improves predictive accuracy but also enhances interpretability by explicitly anchoring predictions to vascular pathology. The multimodal design, therefore, avoids common pitfalls in deep learning models—such as reliance on non-biological artifacts—and instead replicates clinically grounded reasoning patterns used by expert ophthalmologists. 3. Quantitative Results 3.1 Descriptive Statistics and Univariate Analysis The study included 384 preterm infants with complete multimodal data (203 requiring treatment and 181 without treatment; 52.9% vs 47.1%). The mean gestational age was 30.94 ± 3.25 weeks (range: 24–40), the mean birth weight was 1555.41 ± 576.58 g, and the mean NICU stay was 24.30 ± 19.18 days. The majority of infants were male (56.2%), singleton (63.8%), and resided in urban areas (59.9%). ROP stage at the first visit ranged from 0 (38.3%) to 3 (18.0%), with 69.8% of infants showing no plus disease. Univariate analyses revealed significant differences between the treatment and non-treatment groups. Infants requiring treatment had significantly lower gestational age (28.75 ± 2.12 vs 33.40 ± 2.43 weeks; p < 0.001, Cohen’s r = 0.853), lower birth weight (p < 0.001, large effect size), longer NICU stays (p < 0.001), and higher ROP stage and plus disease prevalence (p 0.80). Among the imaging features, PCA1 demonstrated the strongest discriminatory power (r = 0.964, p < 0.001). Demographic variables showed no significant association with treatment requirement. 3.2 Multimodal Model Performance Fourteen machine learning models were evaluated across five predefined multimodal feature sets incorporating PCA embeddings and/or clinical variables. Five-fold stratified cross-validation was performed on the training set (80% of data) to ensure robust performance estimation. Tree-based models (Extra Trees, Random Forest, XGBoost, LightGBM) consistently outperformed other algorithms across feature sets containing PCA embeddings (Table 1 ). For instance, Extra Trees and Random Forest on Set 6 (full multimodal) and Set 4 (clinical + PCA1–3) achieved test accuracy of 0.987, AUC-ROC up to 0.999, and F1-score of 0.988, with near-perfect specificity. Cross-validation confirmed the stability of these results, with mean CV accuracy ranging from 0.984 to 0.986 ± 0.007–0.010 in these sets. In contrast, clinical-only sets (Set 3 and Set 5) exhibited substantially lower performance (maximum test accuracy of 0.896 and 0.948, respectively; CV accuracy 0.888–0.942 ± 0.018–0.026), underscoring the critical contribution of vascular morphology features derived from retinal imaging. Table 1 Performance of top-performing models on multimodal feature sets (test set & 5-fold CV metrics) Feature Set Description Model Accuracy (mean ± std) AUC-ROC (mean ± std) Test Accuracy Test AUC-ROC Test F₁-score Test Sensitivity Test Specificity Set 6 Full multimodal (Clinical/demographic + ROP + PCA1–3) 1. Extra Trees 0.986 ± 0.007 0.995 ± 0.004 0.987 0.999 0.988 0.976 1.000 2. Random Forest 0.985 ± 0.008 0.994 ± 0.005 0.987 0.998 0.988 0.976 1.000 Set 4 Clinical/demographic + PCA1–3 1. Extra Trees 0.984 ± 0.009 0.994 ± 0.005 0.987 0.998 0.988 1.000 0.972 2. Random Forest 0.983 ± 0.010 0.993 ± 0.006 0.987 0.997 0.988 0.976 1.000 Set 2 PCA1–3 + ROP stage/zone/plus 1. Random Forest 0.985 ± 0.008 0.992 ± 0.005 0.987 0.990 0.988 0.976 1.000 2. Extra Trees 0.986 ± 0.007 0.993 ± 0.004 0.987 0.992 0.988 0.976 1.000 Set 5 Clinical/demographic + ROP stage/zone/plus 1. XGBoost 0.942 ± 0.018 0.970 ± 0.012 0.948 0.972 0.952 0.976 0.917 2. Random Forest 0.938 ± 0.020 0.968 ± 0.013 0.935 0.968 0.942 1.000 0.889 Set 3 Clinical/demographic only 1. Random Forest 0.890 ± 0.025 0.935 ± 0.018 0.896 0.938 0.905 0.927 0.861 2. Extra Trees 0.888 ± 0.026 0.933 ± 0.019 0.896 0.935 0.905 0.927 0.861 3.3 Model Interpretability (SHAP Analysis, Feature importance) SHAP analysis was performed on the top-performing models (Extra Trees, Random Forest, XGBoost, and LightGBM) across feature Sets 2, 4, 5, and 6. The results consistently ranked PCA1 as the most influential feature (mean absolute SHAP value > 10–12), followed by ROP stage at the first visit and gestational age. Higher values of PCA1 were strongly associated with an increased likelihood of requiring treatment, reflecting more severe vascular abnormalities in the retina. Synergistic interactions were observed between PCA1 and key clinical risk factors, such as low gestational age, indicating that the models effectively captured the combined effects of imaging-derived vascular morphology and established ROP risk factors. These findings confirm that the models’ predictions are grounded in clinically meaningful features rather than spurious correlations. Discussion The results of this study demonstrate that our explainable multimodal deep learning model achieves high predictive performance for treatment-requiring retinopathy of prematurity, with Extra Trees and Random Forest classifiers yielding a test accuracy of 0.987, AUC-ROC of 0.999, and F1-score of 0.988 on the full multimodal feature set. These outcomes substantially outperform clinical-only models (maximum test accuracy 0.948) and align with or exceed recent AI advancements in ROP prediction, underscoring the value of integrating retinal imaging features—particularly PCA1, which emerged as the most influential predictor in SHAP and Grad-CAM analyses (r = 0.881 with treatment requirement). Our findings are consistent with prior image-only convolutional neural network (CNN) models, which have reported accuracies exceeding 90% for detecting plus disease or treatment-requiring cases (Tan et al. [ 43 ]; Huang et al. [ 44 ]; Yenice et al. [ 45 ]). For instance, Tan et al. achieved 96.6% sensitivity in plus disease detection using DL on fundal images, while Huang et al. reported up to 96.14% sensitivity for detecting ROP (no ROP vs. stages 1–2). However, multimodal approaches combining imaging with clinical data have consistently shown superior performance, particularly in predicting disease progression or reactivation post-treatment (Wu et al. [ 40 ]; He et al. [ 46 ]; Engin et al. [ 47 ]). Wu et al. developed a fusion model with an AUC of 0.822 for ROP reactivation after anti-VEGF, integrating clinical and imaging data, while He et al. achieved an AUC of 0.853 using DNN in the e-ROP study. Our model extends this paradigm by incorporating semi-supervised vessel segmentation via a U-Net-based adversarial domain-adaptation approach, addressing challenges in heterogeneous neonatal fundus images and achieving robust domain-invariant vascular feature extraction—advantages not fully realized in prior studies like Redd et al. [ 48 ], which reported an AUC of 0.960 for detecting type 1 ROP but relied on basic U-Net vessel segmentation without domain adaptation. The attention mechanism in our multi-view fusion further enhances interpretability by adaptively weighting informative retinal perspectives, a critical factor for clinical adoption. The dominance of PCA1 in SHAP analyses highlights the pivotal role of vascular morphology (tortuosity, dilation, and abnormal branching) in identifying severe ROP, consistent with findings from Krishnan et al. [ 49 ], who identified gestational age and device type as key predictors in ML models, and Karkhaneh et al. [ 50 ], who reported 85% sensitivity in digital imaging for referral-warranted ROP in an Iranian cohort. Grad-CAM heatmaps confirmed model attention on the posterior pole, aligning with the International Classification of ROP criteria for plus disease—the strongest predictor of treatment need. This biological alignment addresses a key limitation of many DL models: reliance on non-biological artifacts, as noted in systematic reviews (Jafarizadeh et al.[ 41 ]). Compared to prior multimodal ROP models, our approach offers several advantages. While Takeda et al. [ 51 ] achieved AUCs of 0.93–0.94 using non-imaging ML for ROP occurrence prediction, and Shoeibi et al.[ 52 ] reported 96% accuracy in XGBoost for treatment-needed ROP, these studies often lack external validation, comprehensive clinical integration, or focus on Iranian-specific cohorts (e.g., Karkhaneh et al.). By contrast, our model incorporates a broad range of demographic and neonatal variables (e.g., gestational age, NICU duration) alongside vascular features, achieving near-perfect specificity (up to 1.000) and minimizing false positives—an essential consideration for reducing unnecessary examinations in vulnerable neonates, as emphasized by Engin et al. The use of 5-fold stratified cross-validation on the training set further ensures robustness and generalizability within our cohort, surpassing the internal validation in Wu et al. Despite these strengths, several limitations must be acknowledged. First, the study was conducted at a single tertiary center in northeastern Iran, potentially limiting generalizability to other populations or settings with different ROP prevalence and screening practices, as highlighted in reviews by Jafarizadeh et al. External validation on multicenter or international cohorts—such as those in the e-ROP study (He et al.) or global datasets (Krishnan et al.)—is essential to confirm performance across diverse ethnicities and healthcare systems. Second, while the dataset was rigorously quality-controlled, neonatal fundus imaging remains technically challenging due to motion artifacts and variable illumination; future work could explore additional preprocessing techniques or larger datasets to further enhance robustness. Third, although SHAP and Grad-CAM provided strong interpretability, prospective clinical studies are needed to evaluate the model's real-world impact on screening efficiency and blindness prevention. The clinical implications of this work are significant. Current ROP screening guidelines rely primarily on gestational age and birth weight, resulting in frequent examinations of infants who ultimately do not require treatment. By achieving high accuracy and near-perfect specificity, our model could enable risk stratification, allowing clinicians to prioritize high-risk cases for timely intervention while reducing unnecessary retinal examinations—echoing benefits seen in non-imaging models by Takeda et al. This approach has the potential to decrease infant stress, alleviate parental burden, and optimize resource allocation in neonatal intensive care units, particularly in resource-limited settings like Iran.In conclusion , this explainable multimodal deep learning model represents a promising tool for early prediction of treatment-requiring ROP. By integrating vascular morphology from retinal images with clinical data, the framework not only achieves superior predictive performance but also provides interpretable insights aligned with clinical reasoning. Future multicenter validation and prospective implementation studies, building on global AI advancements (Jafarizadeh et al.), are warranted to translate these findings into routine clinical practice, ultimately contributing to improved outcomes for preterm infants at risk of vision loss. Abbreviations AI: Artificial Intelligence AUC-ROC: Area Under the Receiver Operating Characteristic Curve CNN: Convolutional Neural Network DL: Deep Learning Grad-CAM: Gradient-weighted Class Activation Mapping ICROP: International Classification of Retinopathy of Prematurity ML: Machine Learning NICU: Neonatal Intensive Care Unit PCA: Principal Component Analysis QC: Quality Control ResNet: Residual Network RetCam: RetCam Imaging System RF: Random Forest ROP: Retinopathy of Prematurity SHAP: SHapley Additive exPlanations SVM: Support Vector Machine U-Net: U-shaped Network VEGF: Vascular Endothelial Growth Factor XGBoost: eXtreme Gradient Boosting Declarations Ethics approval and consent to participate This study was approved by the Mashhad University of Medical Sciences Ethics Committee (approval number: IR.MUMS.MEDICAL.REC.1403.338) and adhered to the tenets of the Declaration of Helsinki. Written informed consent was waived due to the retrospective nature of the study and the use of anonymized data. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Funding This study received no specific funding from any funding agency in the public, commercial, or not-for-profit sectors. Author Contribution All authors contributed equally to the study, including conceptualization, data collection, analysis, interpretation, and manuscript preparation. FS, FBA, NS, SE, and AAT each participated in all aspects of the research and approved the final manuscript. Acknowledgements Not applicable. Data Availability The datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request. The complete analysis pipeline code is also available upon request from the corresponding author. References Kong L, Fry M, Al-Samarraie M, Gilbert C, Steinkuller PG. An update on progress and the changing epidemiology of causes of childhood blindness worldwide. J aapos. 2012;16(6):501–7. Du Y, Dang Y. Recent Advances in Understanding the Role of PEST Sequence-Containing Proteins in Retinal Neovascularization. Curr Drug Targets. 2025. Kim SJ, Port AD, Swan R, Campbell JP, Chan RVP, Chiang MF. Retinopathy of prematurity: a review of risk factors and their clinical significance. Surv Ophthalmol. 2018;63(5):618–37. Kaur K, Mikes BA. Retinopathy of Prematurity. StatPearls. Treasure Island (FL): StatPearls Publishing Copyright © 2025. StatPearls Publishing LLC.; 2025. Karkhaneh R, Shokravi N. Assessment of retinopathy of prematurity among 150 premature neonates in Farabi Eye Hospital. Acta Medica Iranica. 2001:35–8. Nakshab M, Bayani GH, Ëshaghi M. Prevalence of retinopathy in premature neonates in neonatal intensive care unit of Boali sina hospital in 2001. J Mazandaran Univ Med Sci. 2003;13(39):63–70. Karkhaneh R, Riazi Esfahani M, Ghojehzadeh L, Kadivar M, Nayeri F, Chams H. Nili Ahmadabadi M. Incidence and risk factors of retinopathy of prematurity. Bina J Ophthalmol. 2005;11(1):81–90. Mansouri MR, Kadivar M, Karkhaneh R, Riazi Esfahani M, Nili Ahmadabadi M, Faghihi H, Mirshahi A, Sadat-Nayeri F, Farahvash MS, Tabatabaei A, Adelpour A. Prevalence and risk factors of retinopathy of prematurity in very low birth weight or low gestational age infants. Bina J Ophthalmol. 2007;12(4):428–34. Naderian G, Iranpour R, Mohammadizadeh M, Fazel Najafabadi F, Badiei Z, Naseri F, et al. The frequency of retinopathy of prematurity in premature infants referred to an ophthalmology clinic in Isfahan. J Isfahan Med Sch. 2011;29(128):126–30. Fayyazi A, Heidarzadeh M, Fayzalahzadeh H, Golzar A, Sadegi K. Prevalence of retinopathy of prematurity in preterm infant hospitalized in Tabriz Alzahra Hospital's NICU. Med J tabriz Univ Med Sci. 2009;30(4):63–6. Sadeghi K, Heidary E, Hashemi F, Heydarzadeh M. Incidence and risk factors of retinopathy of prematurity. Med J Tabriz Univ Med Sci. 2009;30(2):73–7. Riazi-Esfahani M, Alizadeh Y, Karkhaneh R, Mansouri MR, Kadivar M, Ahmadabadi MN, Nayeri F. Retinopathy of prematurity: single versus multiple-birth pregnancies. J ophthalmic Vis Res. 2008;3(1):47. Karkhaneh R, Mousavi SZ, Riazi-Esfahani M, Ebrahimzadeh SA, Roohipoor R, Kadivar M, Ghalichi L, Mohammadi SF, Mansouri MR. Incidence and risk factors of retinopathy of prematurity in a tertiary eye hospital in Tehran. Br J Ophthalmol. 2008;92(11):1446–9. Ghalichi L. Incidence, severity and risk factors for retinopathy of prematurity in premature infants with late retinal examination. Bina J Ophthalmol. 2008 Jul 10. Khatami SF, YOUSEFI AE, FATEHI BG, MAMOURI GA. Retinopathy of prematurity among 1000–2000 gram birth weight newborn infants. Naderian G, Parvaresh M, Rismanchiyan A, Sajadi V. Refractive errors after laser therapy for retinopathy of prematurity. Bina J Ophthalmol. 2009;15(1):13–8. Fouladinejad M, Motahari MM, Gharib MH, Sheishari F, Soltani MO. The prevalence, intensity and some risk factors of retinopathy of premature newborns in Taleghani Hospital, Gorgan, Iran. Saeidi R, Hashemzadeh A, Ahmadi S. RAHMANI S. Prevalence and predisposing factors of retinopathy of prematurity in very low-birth-weight infants discharged from NICU. Mousavi SZ, Karkhaneh R, Riazi-Esfahani M, Mansouri MR, Roohipoor R, Ghalichi L, Kadivar M, Nili-Ahmadabadi M, Naieri F. Retinopathy of prematurity in infants with late retinal examination. J ophthalmic Vis Res. 2009;4(1):24. NADERIAN GA, MOULAVI VH, Hadipour M, Sajadi V. Prevalence and risk factors for retinopathy of prematurity in Isfahan. Bayat-Mokhtari M, Pishva N, Attarzadeh A, Hosseini H, Pourarian S. Incidence and risk factors of retinopathy of prematurity among preterm infants in Shiraz/Iran. Iran J Pediatr. 2010;20(3):303. Ebrahim M, Ahmad RS, Mohammad M. Incidence and risk factors of retinopathy of prematurity in Babol, North of Iran. Ophthalmic Epidemiol. 2010;17(3):166–70. MOUSAVI SZ, RIAZI EM, Rouhipour R, JABARVAND M, Ghalichi L, NILI AM, GHASEMI F, AALAMI HZ, Karkhaneh R. Characteristics of advanced stages of retinopathy of prematurity. Ghaseminejad A, Niknafs P. Distribution of retinopathy of prematurity and its risk factors. Iran J Pediatr. 2011;21(2):209. Afarid M, Hosseini H, Abtahi B. Screening for retinopathy of prematurity in South of Iran. Middle East Afr J Ophthalmol. 2012;19(3):277–81. Feghhi M, Altayeb SM, Haghi F, Kasiri A, Farahi F, Dehdashtyan M, Movasaghi M, Rahim F. Incidence of retinopathy of prematurity and risk factors in the south-western region of Iran. Middle East Afr J Ophthalmol. 2012;19(1):101–6. Abrishami M, Maemori GA, Boskabadi H, Yaeghobi Z, Mafi-Nejad S, Abrishami M. Incidence and risk factors of retinopathy of prematurity in mashhad, northeast iran. Iran Red Crescent Med J. 2013;15(3):229. Kazem Sabzehei M, Afje h SA, Farahani AD, Shamshiri AR, Esmaili F. Retinopathy of Prematurity: Incidence, Risk Factors, and Outcome. Archives Iran Med (AIM). 2013;16(9). Ahmadpour-kacho M, Pasha YZ, Rasoulinejad SA, Hajiahmadi M, Pourdad P. Correlation between retinopathy of prematurity and clinical risk index for babies score. Tehran Univ Med J. 2014;72(6). Rasoulinejad SA, Montazeri M. Retinopathy of prematurity in neonates and its risk factors: a seven year study in northern Iran. Open Ophthalmol J. 2016;10:17. Broxterman EC, Hug DA. Retinopathy of Prematurity: A Review of Current Screening Guidelines and Treatment Options. Mo Med. 2016;113(3):187–90. Kaur K, Gurnani B, Kannusamy V, Yadalla D. A tale of orbital cellulitis and retinopathy of prematurity in an infant: First case report. Eur J Ophthalmol. 2022;32(6):Np20–3. Hutchinson AK, Melia M, Yang MB, VanderVeen DK, Wilson LB, Lambert SR. Clinical Models and Algorithms for the Prediction of Retinopathy of Prematurity: A Report by the American Academy of Ophthalmology. Ophthalmology. 2016;123(4):804–16. Palmer EA, Flynn JT, Hardy RJ, Phelps DL, Phillips CL, Schaffer DB, et al. Incidence and early course of retinopathy of prematurity. The Cryotherapy for Retinopathy of Prematurity Cooperative Group. Ophthalmology. 1991;98(11):1628–40. Revised indications for the. treatment of retinopathy of prematurity: results of the early treatment for retinopathy of prematurity randomized trial. Arch Ophthalmol. 2003;121(12):1684–94. Ophthalmology So P, AAo O, AAo. Ophthalmology AAfP, Strabismus. Screening Examination of Premature Infants for Retinopathy of Prematurity. Pediatrics. 2006;117(2):572–6. Desai S, Athikarisamy SE, Lundgren P, Simmer K, Lam GC. Validation of WINROP (online prediction model) to identify severe retinopathy of prematurity (ROP) in an Australian preterm population: a retrospective study. Eye (Lond). 2021;35(5):1334–9. Shoeibi N, Ameri N, Hoseinkhani MR, Ansari Astaneh MR, Motamed Shariati M, Hosseini SM, Abrishami M, Abrishami M, Zamani G, Heidarzadeh HR. Deep learning algorithms for timely diagnosis of retinopathy of prematurity requiring treatment. Eye 2025 Nov 3:1–8. Wang J, Ji J, Zhang M, Lin JW, Zhang G, Gong W, Cen LP, Lu Y, Huang X, Huang D, Li T. Automated explainable multidimensional deep learning platform of retinal images for retinopathy of prematurity screening. JAMA Netw open. 2021;4(5):e218758. Wu R, Zhang Y, Huang P, Xie Y, Wang J, Wang S, Lin Q, Bai Y, Feng S, Cai N, Lu X. Prediction of Reactivation After Antivascular Endothelial Growth Factor Monotherapy for Retinopathy of Prematurity: Multimodal Machine Learning Model Study. J Med Internet Res. 2025;27:e60367. Jafarizadeh A, Maleki SF, Pouya P, Sobhi N, Abdollahi M, Pedrammehr S, Lim CP, Asadi H, Alizadehsani R, Tan RS, Islam SM. Current and future roles of artificial intelligence in retinopathy of prematurity. Artif Intell Rev. 2025;58(6):1–55. Zhang Y, Yang L, Chen J, Fredericksen M, Hughes DP, Chen DZ. Deep adversarial networks for biomedical image segmentation utilizing unannotated images. InInternational conference on medical image computing and computer-assisted intervention. 2017 Sep 4 (pp. 408–416). Cham: Springer International Publishing. Tαν Z, Siµkiν S, Lαi C, Dαi S. Deeπ leαρνiνg αlgoρiτhµ φoρ αuτoµατeδ δiαgνoσiσ oφ ρeτiνoπατhy oφ πρeµατuρiτy πluσ δiσeασe. Tρανσlατioναl viσioν σcieνce τechνology. 2019;8(6):23. Huang YP, Basanta H, Kang EY, Chen KJ, Hwang YS, Lai CC, Campbell JP, Chiang MF, Chan RV, Kusaka S, Fukushima Y. Automated detection of early-stage ROP using a deep convolutional neural network. Br J Ophthalmol. 2021;105(8):1099–103. Kıran Yenice E, Kara C, Erdaş ÇB. Automated detection of type 1 ROP, type 2 ROP and A-ROP based on deep learning. Eye. 2024;38(13):2644–8. He D, Luo X, Ying B, Quinn GE, Baumritter A, Chen Y, Ying GS, He L. Machine learning models for predicting treatment-requiring retinopathy of prematurity in the e-ROP study. Translational Vis Sci Technol. 2025;14(8):14. Durmaz Engin C, Ozturk T, Ozkan O, Oztas A, Selver MA, Tuzun F. Prediction of retinopathy of prematurity development and treatment need with machine learning models. BMC Ophthalmol. 2025;25(1):1–2. Redd TK, Campbell JP, Brown JM, Kim SJ, Ostmo S, Chan RVP, et al. Evaluation of a deep learning image assessment system for detecting severe retinopathy of prematurity. Br J Ophthalmol. 2019;103(5):580–4. Kρiσhναν A, Sαlµαν Sαρkαρ M, Bαδαvατh L. Mαchiνe Leαρνiνg Aππlicατioνσ iν Reτiνoπατhy oφ Pρeµατuρiτy Diαgνoσiσ Uσiνg τhe ROP Reτiναl Iµαge Dατασeτ. J Neoναταl Suρg [Iντeρνeτ]. 2025 Φeβ. 7 [ciτeδ 2025 Dec. 28];14(1S):820-7. Avαilαβle φρoµ: hττπσ://jνeoναταlσuρg.coµ/iνδex.πhπ/jνσ/αρτicle/vieω/1607 Kαρkhανeh R, Ahµαδραji A, Eσφαhανi MR, Roohiπouρ R, Dαστjανi AΦ, Iµανi M, Khoδαβανδe A, Eβραhiµiαδiβ N, Ahµαδαβαδi MN. The αccuραcy oφ δigiταl iµαgiνg iν δiαgνoσiσ oφ ρeτiνoπατhy oφ πρeµατuρiτy iν Iραν: α πiloτ στuδy. Jouρναl oφ Oπhτhαlµic & Viσioν Reσeαρch. 2019 Jαν;14(1):38. Takeda Y, Kaneko Y, Sugimoto M, Yamashita H, Sasaki A, Mitsui T. Prediction Models for Retinopathy of Prematurity Using Nonimaging Machine Learning Approaches: A Regional Multicenter Study. Ophthalmol Sci. 2025;5(4):100715. Shoeibi N, Abrishami M, Hosseini SM, Ansari-Astaneh MR, Farrahi R, Gharib B, et al. Development and validation of machine learning classifiers for predicting treatment-needed retinopathy of prematurity. BMC Med Inf Decis Mak. 2025;25(1):221. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 26 Mar, 2026 Reviews received at journal 25 Mar, 2026 Reviews received at journal 16 Mar, 2026 Reviews received at journal 15 Mar, 2026 Reviewers agreed at journal 15 Mar, 2026 Reviews received at journal 13 Mar, 2026 Reviewers agreed at journal 12 Mar, 2026 Reviewers agreed at journal 11 Mar, 2026 Reviewers agreed at journal 09 Mar, 2026 Reviewers invited by journal 09 Mar, 2026 Editor invited by journal 12 Feb, 2026 Editor assigned by journal 11 Feb, 2026 Submission checks completed at journal 11 Feb, 2026 First submitted to journal 10 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8838478","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":604382475,"identity":"8098193d-acd8-40d8-8506-99c06facd095","order_by":0,"name":"Fatemeh Sefidbaf","email":"","orcid":"","institution":"Mashhad University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Fatemeh","middleName":"","lastName":"Sefidbaf","suffix":""},{"id":604382476,"identity":"bba79b93-0f5f-44e7-bb3f-bd1b13bc9dc5","order_by":1,"name":"Fatemeh Baharvand Ahmadi","email":"","orcid":"","institution":"Mashhad University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Fatemeh","middleName":"Baharvand","lastName":"Ahmadi","suffix":""},{"id":604382477,"identity":"d9bc0138-02b7-4969-8f2b-2ecf07254cce","order_by":2,"name":"Nasser Shoeibi","email":"","orcid":"","institution":"Mashhad University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Nasser","middleName":"","lastName":"Shoeibi","suffix":""},{"id":604382478,"identity":"7c2bb1f1-eb33-4194-abda-be833ac155fe","order_by":3,"name":"Saeid Eslami","email":"","orcid":"","institution":"Mashhad University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Saeid","middleName":"","lastName":"Eslami","suffix":""},{"id":604382479,"identity":"fdeeff34-5748-4bf9-b216-0757f94a73a6","order_by":4,"name":"Amin Amiri Tehranizadeh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAu0lEQVRIiWNgGAWjYBACNgbGBobENgk5JBECgB+s5ZyNMfFaJBuABOO/tMQGoh1mcO1w24eH2w6nr52RnfyBocaOgU/6AAEttxObZyRuO5y77UbuNgmGY8kMbHwJhLUwwLQAPXKAgY2HgMPswVra/qeb3cjd/IHhHxFaILa0HU4AatkgwdhGghbDbWfebpNI7EvmIUJL+mPGn22H5c2OAx324ZudnHwPAS2oIIGBgZAdo2AUjIJRMAqIAQB+OUa5e5g1+QAAAABJRU5ErkJggg==","orcid":"","institution":"Mashhad University of Medical Sciences","correspondingAuthor":true,"prefix":"","firstName":"Amin","middleName":"Amiri","lastName":"Tehranizadeh","suffix":""}],"badges":[],"createdAt":"2026-02-10 08:38:59","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8838478/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8838478/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104554808,"identity":"85dc0ab6-6abd-42a4-9686-377215930783","added_by":"auto","created_at":"2026-03-13 08:58:09","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":501331,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the proposed semi-supervised learning framework for ROP analysis. This pipeline leverages segmentation, reconstruction, and adversarial consistency to learn vessel-aware features for predicting treatment-requiring ROP. Training proceeds in three stages: supervised segmentation, semi-supervised representation learning with frozen segmentation weights, and classifier training using the learned features while the reconstruction branch remains fixed. Green arrows indicate trainable components, red arrows denote frozen modules, and colored frames represent target-domain labeled (red) and source-domain unlabeled (blue) images.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/a2f25ff7e57ae5358f1038c7.jpeg"},{"id":104554820,"identity":"fa5ed7cc-4c44-4388-a727-ba39ee790a05","added_by":"auto","created_at":"2026-03-13 08:58:13","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":291194,"visible":true,"origin":"","legend":"\u003cp\u003eDeep learning model development pipeline: The framework is built in two stages: (1) training a baseline single-image classifier using an ImageNet-pretrained encoder, and (2) extending the model to a multi-view architecture that processes four directional fundus images and their corresponding vessel segmentation maps through parallel pretrained encoders. Single-channel segmentation inputs are integrated by adapting the first convolutional layer via channel-wise kernel averaging, enabling consistent vessel-aware feature extraction across views.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/3429f33903d7769ae31bab14.jpeg"},{"id":104554786,"identity":"71c138c3-9a80-4888-80a7-c899423f5390","added_by":"auto","created_at":"2026-03-13 08:58:05","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1954291,"visible":true,"origin":"","legend":"\u003cp\u003eExamples of treatment-requiring ROP cases shown across four fundus views. Each column represents one of the four directional fundus views. The rows, from top to bottom, correspond to the original fundus images, the unsupervised segmentation outputs, and the semi-supervised segmentation results.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/04778dbc0c8dd1c053979de3.jpeg"},{"id":104554858,"identity":"f9f84679-9a39-4d27-9880-d28bce40c327","added_by":"auto","created_at":"2026-03-13 08:58:14","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":2263913,"visible":true,"origin":"","legend":"\u003cp\u003eExamples of non–treatment-requiring ROP cases in four views. Each column represents one of the four directional fundus views. The rows, from top to bottom, correspond to the original fundus images, the unsupervised segmentation outputs, and the semi-supervised segmentation results.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/df77c1c63b954d5733933f26.jpeg"},{"id":104554785,"identity":"ab37673b-0f7f-432a-afe9-066df2a224ab","added_by":"auto","created_at":"2026-03-13 08:58:05","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":2326086,"visible":true,"origin":"","legend":"\u003cp\u003eModel interpretability analysis. Interpretability heatmaps highlight the retinal regions most influential to the model’s predictions.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/a5b8fa537fe04fbf9d91cf0b.jpeg"},{"id":104554861,"identity":"6bb42cb6-20aa-4644-8482-766d2c888aec","added_by":"auto","created_at":"2026-03-13 08:58:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":8226491,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8838478/v1/bbc5b603-2ff1-4b96-b92c-caaf0ac83afc.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Explainable Multimodal Deep Learning Model for Early Prediction of Treatment-Requiring Retinopathy of Prematurity","fulltext":[{"header":"Introduction","content":"\u003cp\u003eRetinopathy of Prematurity (ROP) is a disease characterized by abnormal retinal blood vessel growth in preterm newborns [\u003cspan class=\"CitationRef\"\u003e1\u003c/span\u003e]. The condition typically starts when these infants are exposed to high oxygen levels (hyperoxia), causing existing vessels to close and regress. When oxygen levels normalize again, the resulting relative hypoxia drives excessive new vessel formation, mainly through vascular endothelial growth factor (VEGF) [\u003cspan class=\"CitationRef\"\u003e2\u003c/span\u003e]. Known risk factors include extremely low gestational age or birth weight, inflammation during pregnancy or after birth, the general immaturity of premature infants, lung-related complications, and anemia [\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e]. Each year, around 15\u0026nbsp;million premature infants are born globally, placing approximately 1.4\u0026nbsp;million at risk for ROP [\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e]. Roughly 32,000 of these infants progress to severe visual impairment or blindness [\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e]. In Iran, prevalence rates reported in various studies range widely from 5.6% to 70.3% [\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e–\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e]. Available treatments include laser photocoagulation, cryotherapy, and anti-VEGF injections. Still, early diagnosis via timely screening is essential for successful management [\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e]. Current guidelines rely mostly on gestational age and birth weight, which frequently leads to screening many infants who will not develop treatment-requiring disease [\u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e–\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e]. Frequent eye examinations can be stressful and uncomfortable for the infants, while also creating emotional and financial strain for families [\u003cspan class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e37\u003c/span\u003e]. Factors like delayed referrals, missed appointments, shortage of pediatric retinal specialists, and inadequate parental awareness further raise the chances of preventable blindness [\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e]. Recent studies have shown accuracies over 90% for image-only models in detecting plus disease and cases that need intervention. These results come from applying convolutional neural networks to fundus photographs (Shoeibi et al., 2025[\u003cspan class=\"CitationRef\"\u003e38\u003c/span\u003e]; Wang et al., 2021[\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e]). Multimodal approaches – the ones that mix retinal images with clinical and demographic data – usually do better than image-only models. This advantage is clearest in predicting disease progression or reactivation after anti-VEGF treatment (Wu et al., 2025[\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e]). Researchers have also come up with explainable AI methods. These provide clearer interpretations of ROP stages and key pathological features, which helps increase trust among clinicians and makes adoption easier (Wang et al., 2021[\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e]). On the other hand, systematic reviews make it clear that many of these models still have drawbacks: not enough external validation, datasets that are too small, and limited inclusion of wider clinical factors. All this points to a continuing need for stronger, more widely applicable, and truly interpretable multimodal systems (Jafarizadeh et al., 2025[\u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e]). The present study introduces an explainable multimodal deep learning model. It combines retinal images with a broad range of clinical and demographic variables – things like infant sex, parental education, pregnancy and birth characteristics, neonatal factors (including NICU stay duration), and eye-specific severity measures – to predict treatment-requiring ROP. The model is designed to enhance risk stratification, allow earlier detection of high-risk cases, support timely treatment, minimize unnecessary examinations, reduce healthcare system burden, and ultimately improve care quality and blindness prevention in this vulnerable population.\u003c/p\u003e \n\n\n\n \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003c/div\u003e \u003c/div\u003e\n\n \n\n \n\n"},{"header":"Methods","content":"\u003ch3\u003e1. Study Design and Participant Cohort\u003c/h3\u003e\u003cp\u003e This retrospective cohort study was conducted at Khatam-al-Anbia Ophthalmology Hospital, the primary tertiary referral center for neonatal ophthalmic care in northeastern Iran. The study was approved by the Mashhad University of Medical Sciences Ethics Committee (IR.MUMS.MEDICAL.REC.1403.338) and adhered to the tenets of the Declaration of Helsinki. Infants were systematically referred for retinopathy of prematurity screening according to national guidelines, which specify eligibility as gestational age \u0026lt; 34 weeks and/or birth weight ≤ 2000 g. Consecutive infants presenting to the dedicated ROP clinic between January 2021 and December 2024 were included. All infants underwent standardized comprehensive ROP evaluation using the RetCam imaging system, documenting ROP stage, zone, and presence of plus disease. Infants with complete evaluation data were eligible; those with unresolved missing data on key variables were excluded. Treatment eligibility (laser photocoagulation or intravitreal anti-VEGF injection) was independently assessed by two senior pediatric ophthalmologists with \u0026gt; 10 years of experience in ROP management. Discrepancies were adjudicated by a third senior ophthalmologist, yielding a kappa coefficient of 0.89 (95% CI 0.85–0.93). From the screened population, 400 treatment-requiring infants were identified. A comparison group of 400 infants without treatment indication was randomly selected. Following data quality assessment, the final analytical cohort comprised 384 infants (203 treated, 181 untreated) with complete multimodal datasets. Written informed consent was waived by the ethics committee due to the retrospective design and use of fully anonymized data.\u003c/p\u003e\u003ch3\u003e2. Data Acquisition and Preprocessing\u003c/h3\u003e\u003ch2\u003e2.1. Clinical and Demographic Variables\u003c/h2\u003e\u003cp\u003eClinical and demographic data were extracted from electronic health records using Python 3.12 with pandas. Variables included infant sex, twin status, ROP features at first visit (stage, zone, and plus disease), gestational age and birth weight, age and weight at first ROP visit, NICU stay duration, parental ages at birth, parental education levels, gravidity, pregnancy initiation mode, pre-pregnancy maternal status, and city of residence. After data cleaning, no missing values remained in the analyzed variables. Numerical variables underwent outlier removal (interquartile range method) and standardization (using StandardScaler fitted on the training set). Categorical variables were label-encoded (using LabelEncoder fitted on the training set). The target variable was binary treatment requirement (1 = treatment required, 0 = no treatment).\u003c/p\u003e\u003ch2\u003e2.2. ROP image selection and preprocessing\u003c/h2\u003e\u003ch2\u003e2.2.1. Image Quality Control\u003c/h2\u003e\u003cp\u003eAll retinal images underwent a rigorous quality control (QC) procedure prior to inclusion in the analytical pipeline to ensure reliability of downstream segmentation, feature extraction, and classification tasks. The heterogeneous nature of neonatal fundus imaging—characterized by variable illumination, motion artifacts, media opacity, and differences in field-of-view—necessitates stringent QC to prevent poor-quality images from degrading model performance or introducing bias. Therefore, a standardized multi-stage QC framework was implemented to systematically evaluate the suitability of each image. First, all fundus photographs were screened by trained graders for essential diagnostic visibility, including adequate depiction of the posterior pole, optic disc, and major vascular arcades. Images with severe blur, defocus, or motion streaks that obscured vascular structures were excluded. Particular attention was given to vessel clarity and contrast, as vessel segmentation accuracy and subsequent vascular feature extraction depend heavily on the visibility of fine vascular structures. Images with excessive glare, overexposure, or shadowing caused by eyelids or specular reflections were removed if these artifacts interfered with vessel morphology assessment. Cases with incomplete field-of-view that did not capture the posterior pole or displayed only peripheral retina were also excluded, as these images do not provide sufficient information for evaluating plus disease or predicting treatment-requiring ROP. Following the manual grading step, automated quality metrics were applied to ensure consistency across the dataset. These included evaluations of illumination uniformity, contrast distribution, edge sharpness, and signal-to-noise ratio. Images failing predefined thresholds were flagged for secondary review and excluded when quality was deemed insufficient for reliable vessel segmentation. After completing all QC stages, the four highest-quality fundus images for each eye—captured from four standard directional views—were selected as the final candidates for subsequent processing phases, ensuring that downstream algorithms operated on the most informative and diagnostically reliable representations of the retina. This hybrid QC approach—combining expert review with automated assessment—ensured that only images meeting minimum diagnostic and computational quality requirements were used for model development.\u003c/p\u003e\u003ch2\u003e2.2.2 Semi-supervised Vessel Segmentation\u003c/h2\u003e\u003cp\u003eGiven the scarcity of expertly annotated retinal vessel masks and the abundance of unlabeled ROP fundus images typically encountered in clinical settings, we adopted a semi-supervised segmentation strategy inspired by the Deep Adversarial Network framework of Zhang et al. [\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e]. As illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, the model employs a U-Net architecture with a ResNet-34 encoder pretrained on ImageNet to enhance feature generalization and training stability. In the supervised warm-up phase, the network was trained using high-quality labeled retinal vessel datasets (DRIVE, STARE, CHASEDB1), with a combined Dice and binary cross-entropy loss.\u003c/p\u003e\u003cp\u003eIn the second stage, a dual-branch encoder–decoder model was introduced to jointly process both labeled target-domain images and unlabeled source-domain images. The public datasets were defined as the target domain, representing high-quality labeled vascular structures, whereas the locally acquired ROP images constituted the source domain, reflecting real-world clinical variability. During this phase, the segmentation branch of the network was frozen to preserve the vascular knowledge learned from the target domain, while the reconstruction branch remained trainable to encourage the extraction of domain-invariant retinal representations. Adversarial and reconstruction-based consistency constraints were applied to ensure that the feature distributions of source-domain images gradually aligned with those of the target domain. Importantly, the classifier associated with this stage was trained to output True for target-domain images and False for source-domain images, enabling the model to explicitly learn domain discrimination. This domain-classification signal plays a key role in guiding the encoder toward representations that minimize domain shift, thereby improving robustness in subsequent clinical prediction tasks. In the third stage of the framework, the dual-branch encoder–decoder architecture operates under an adversarial domain-alignment objective. At this point, the reconstruction branch of the network is frozen, preserving the domain-invariant structural representations learned in previous phases, while the segmentation branch remains fully trainable to allow targeted refinement of vessel probability maps for source-domain ROP images. Simultaneously, the domain-classification module of the framework is also frozen, ensuring that its decision boundary remains fixed during this phase. The segmentation branch is then optimized such that its outputs for source-domain images are increasingly likely to be interpreted by the frozen classifier as target-domain segmentation maps. This setup forces the segmentation network to adapt its output distribution toward that of the labeled target-domain datasets, thereby reducing residual domain discrepancies. As a result, the segmentation model learns to generate vessel maps for heterogeneous, real-world ROP images that closely resemble the structured vascular patterns seen in high-quality public datasets, achieving robust domain-invariant segmentation performance.\u003c/p\u003e\u003cp\u003e \u003c/p\u003e\u003ch3\u003e3. DL model development\u003c/h3\u003e\u003cp\u003eAs illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e, The development of the proposed deep learning framework proceeded in two main stages, beginning with a single-image classification model and subsequently extending to a multi-view architecture capable of leveraging four directional fundus images from each eye. The goal of this staged development was to establish a robust baseline classifier and then enhance it through the incorporation of semi-supervised segmentation priors and multi-task feature extraction.\u003c/p\u003e\u003cp\u003eIn the first stage, a standard image-based classifier was trained to predict treatment-requiring ROP from a single posterior-pole fundus image. A convolutional neural network pretrained on ImageNet served as the encoder backbone, enabling the model to benefit from strong visual feature representations despite limited domain-specific data. The classifier was optimized end-to-end using cross-entropy loss, and this single-image model provided an essential baseline for understanding the discriminative retinal features associated with clinically significant ROP.\u003c/p\u003e\u003cp\u003eIn the second stage, the framework was expanded to incorporate all four available fundus images captured from different viewpoints around the optic disc. For each view, vessel segmentation probability maps generated from the preceding semi-supervised segmentation model were used as additional structural inputs. Each fundus image and its corresponding vessel map were passed through one of four encoder streams, all sharing the same architecture but initialized with pretrained weights from the earlier training phase. To accommodate the single-channel vessel segmentation maps while preserving the pretrained structure of the encoder, the first convolutional layer was adapted by computing the channel-wise mean of its pretrained ImageNet filters and applying this averaged kernel to the segmentation input; after this modification of the initial layer, all subsequent convolutional blocks remained identical to the pretrained backbone. Thus, the fundus images were processed using the original three-channel pretrained weights, whereas the segmentation masks were integrated through a minimally adjusted first layer followed by the same deeper encoder architecture.\u003c/p\u003e\u003cp\u003e \u003c/p\u003e\u003ch3\u003e4. Multimodal Fusion and Machine Learning Evaluation\u003c/h3\u003e\u003cp\u003eThe multimodal fusion framework integrates four directional fundus images and their corresponding vessel segmentation maps with structured demographic and clinical information to generate comprehensive representations for predicting treatment-requiring ROP. Each fundus image is paired with its vessel probability map and processed jointly through one of four parallel encoder streams with shared architecture. These encoders are initialized with weights pretrained in earlier model development stages, allowing extraction of stable, vessel-aware feature embeddings across retinal viewpoints. Independent processing of each directional view captures spatially distributed vascular cues—including tortuosity, dilation, branching complexity, and disc-centered geometry—that vary with camera angle and illumination conditions. The four feature vectors are projected into a shared latent space and fused using a multi-head attention module that adaptively weights the contribution of each view, prioritizing informative perspectives while suppressing noisier inputs. This produces a robust fused imaging representation. PCA dimensionality reduction is applied to this representation to generate low-dimensional embeddings (PCA1–3).\u003c/p\u003e\u003cp\u003eAll clinical and demographic variables included in the study, as well as the composition of the multimodal feature sets, were selected based on clinical expertise from experienced pediatric ophthalmologists. Selection prioritized variables with established or potential associations with ROP severity and treatment requirement.\u003c/p\u003e\u003cp\u003eThe fused imaging representation (including PCA1–3) is concatenated with clinical/demographic variables to form five predefined multimodal feature sets for comparative evaluation:\u003c/p\u003e\u003cp\u003e \u003c/p\u003e\u003cul\u003e \u003cli\u003e \u003cp\u003eSet 1: PCA1–3, ROP stage, zone, plus disease at first visit.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSet 2: Clinical/demographic variables (birth age/weight, age/weight at first visit, NICU duration, parental age/education, sex, twin status, pregnancy initiation, pre-pregnancy issues, gravidity, city of residence).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSet 3: Set 2 + PCA1–3.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSet 4: Set 2 + ROP stage, zone, plus disease at first visit.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSet 5: Set 2 + PCA1–3 + ROP stage, zone, plus disease at first visit.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSet 6: All features (Set 1 + Set 2)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eData were split 80/20 (stratified by target, random_state = 42). Preprocessing was fitted on the training set only.\u003c/p\u003e\u003cp\u003eGiven the mild class imbalance in the dataset (Class Ratio 0:1 = 0.89:1, corresponding to approximately 89 non-treatment cases per treatment case), class weighting was applied during model training to mitigate bias toward the majority class without altering the original data distribution. Specifically, the class_weight=balanced parameter was used in tree-based ensembles (Random Forest, Extra Trees, XGBoost, Gradient Boosting, AdaBoost, Decision Tree) and equivalent weighting strategies were implemented in other classifiers (e.g., class_weight=balanced in Logistic Regression and SVM). This approach was preferred over resampling techniques (e.g., oversampling or undersampling) due to its computational efficiency, preservation of the original data distribution, and avoidance of potential artifacts from synthetic sample generation in a clinically sensitive and moderately sized dataset. The following machine learning models were evaluated on each feature set: Random Forest (n_estimators = 200, max_depth = 10, class_weight='balanced'), XGBoost (n_estimators = 150, max_depth = 6, learning_rate = 0.1), Gradient Boosting, Logistic Regression (C = 0.1, penalty='l2'), SVM (kernel='rbf', C = 1.0, gamma='scale'), MLP (hidden_layers=(50,25), alpha = 0.0001), AdaBoost, Extra Trees, Decision Tree, KNN (n_neighbors = 5), Gaussian Naive Bayes, LDA, and QDA. By combining multi-view vascular morphology with clinical context across different feature sets, the framework enables comparative assessment of multimodal contributions to treatment prediction.\u003c/p\u003e\u003ch3\u003e5. Statistical Analysis of Feature Importance\u003c/h3\u003e\u003cp\u003eUnivariate statistical analyses were performed to assess the importance and association of individual features with treatment requirement. Descriptive statistics (mean, SD, median, min/max, skewness, kurtosis for numerical; frequency/percentage for categorical) were computed. Normality was tested using Shapiro-Wilk/Kolmogorov-Smirnov. Group comparisons used automated test selection: Student's t-test/Welch's t-test for normal data, Mann-Whitney U for non-normal. Categorical associations used χ² or Fisher's exact test with Cramér's V. Pearson correlation was calculated for numerical variables with treatment. These statistical tests provided evidence of feature importance without formal automated selection.\u003c/p\u003e"},{"header":"Results","content":"\n\u003ch3\u003e1. Semi-supervised Vessel Segmentation\u003c/h3\u003e\n\u003cp\u003eFigures \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e present qualitative comparisons of retinal vessel segmentation outputs generated by the proposed semi-supervised framework across multiple ROP cases, each containing four directional fundus views. For each sample, the original ROP images are shown in the top row, followed by the vessel maps produced using an unsupervised baseline method and the refined vessel masks obtained through the proposed semi-supervised segmentation model. These visual results demonstrate the substantial improvements achieved when unlabeled ROP images are leveraged through adversarial and reconstruction-based consistency training. Across all samples, the unsupervised segmentation outputs exhibit common failure patterns, including fragmented vessel structures, loss of fine-caliber peripheral vessels, and incomplete delineation near the optic disc region. These limitations are particularly evident in fundus views with uneven illumination, motion blur, or glare\u0026mdash;conditions frequently encountered in neonatal imaging. Additionally, unsupervised masks often show excessive noise or spurious edges in low-contrast regions, reducing their suitability for downstream vascular feature extraction. By contrast, the semi-supervised segmentation maps display markedly enhanced anatomical coherence and vessel continuity across all four views. The proposed method successfully recovers subtle vascular branches, improves delineation of tortuous vessels, and reduces background noise, even in images with challenging acquisition conditions. The integration of labeled public fundus datasets during supervised warm-up, followed by domain adaptation on unlabeled ROP images, enables the model to generalize effectively to neonatal retinal morphology, which differs substantially from adult fundus patterns. This is particularly notable in peripheral zones where vessel caliber is small and prone to segmentation errors in conventional unsupervised methods.Moreover, the semi-supervised approach yields consistent vessel structures across all views within each sample, demonstrating the model\u0026rsquo;s robustness to variations in angle, illumination, and retinal curvature. This cross-view consistency is essential because the segmentation outputs are subsequently used as structural priors in the multi-view classification framework. Accurate preservation of vascular geometry, especially dilation, branching density, and tortuosity is critical for downstream prediction of treatment-requiring ROP. The improved vessel fidelity achieved via semi-supervised segmentation substantially enhances the reliability of extracted vascular biomarkers.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003e2. Model Interpretability Analysis\u003c/h3\u003e\n\u003cp\u003eTo evaluate the clinical validity and transparency of the proposed predictive framework, we performed an extensive interpretability analysis using gradient-based class activation mapping (Grad-CAM) applied to the final multimodal fusion model. The goal of this analysis was to determine whether the model relies on physiologically meaningful retinal features\u0026mdash;particularly vascular abnormalities, when predicting treatment-requiring ROP.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e presents representative examples of the original fundus images, their corresponding vessel segmentation maps, and the model-generated attention heatmaps. Across all samples, the visual explanations consistently demonstrated a strong alignment between the model\u0026rsquo;s decision-making process and established clinical hallmarks of severe ROP. In nearly all cases, the Grad-CAM activation maps concentrated on the posterior pole, including the optic disc and major vascular arcades. This corresponds precisely to the anatomical region where plus disease\u0026mdash;the strongest predictor of treatment requirement\u0026mdash;is assessed by clinicians. Increased tortuosity and dilation of central vessels are central diagnostic indicators in the ICROP classification system, and the model\u0026rsquo;s selective attention to these regions suggests that it has successfully learned vascular patterns associated with advanced disease. Notably, the heatmaps rarely emphasized peripheral retinal regions, which aligns with clinical practice, as treatment decisions are largely driven by posterior vascular abnormalities rather than the peripheral stage alone. This focal attention on the posterior pole provides strong evidence that the model\u0026rsquo;s predictions are grounded in pathophysiologic cues rather than spurious correlations. Overlaying the heatmaps onto the segmentation maps further confirmed that the model\u0026rsquo;s attention aligns with specific vascular structures rather than non-informative image regions such as illumination artifacts, imaging borders, or peripheral blur. In multiple examples, activation clusters followed the trajectory of dilated or tortuous vessels, indicating that the model leverages vessel morphology\u0026mdash;rather than global image color or brightness\u0026mdash;to guide its predictions. In cases with highly engorged vessels or asymmetric vascular branching, the heatmaps highlighted exactly those aberrant areas, demonstrating sensitivity to clinically meaningful nuances. For images with milder vascular changes, the attention maps were more diffuse yet still centered around the arcades, reflecting the model\u0026rsquo;s uncertainty in borderline cases\u0026mdash;behavior that mimics real-world variability among human graders. Importantly, the interpretability results reveal consistency between the vessel-aware representations learned during the semi-supervised segmentation stage and the downstream classification behavior. Because the fusion model integrates both raw fundus features and segmentation-derived vascular context, the resulting attention maps capture a blend of structural and morphological cues. This demonstrates that the segmentation-guided feature extraction process not only improves predictive accuracy but also enhances interpretability by explicitly anchoring predictions to vascular pathology. The multimodal design, therefore, avoids common pitfalls in deep learning models\u0026mdash;such as reliance on non-biological artifacts\u0026mdash;and instead replicates clinically grounded reasoning patterns used by expert ophthalmologists.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003e3. Quantitative Results\u003c/h3\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Descriptive Statistics and Univariate Analysis\u003c/h2\u003e \u003cp\u003eThe study included 384 preterm infants with complete multimodal data (203 requiring treatment and 181 without treatment; 52.9% vs 47.1%). The mean gestational age was 30.94\u0026thinsp;\u0026plusmn;\u0026thinsp;3.25 weeks (range: 24\u0026ndash;40), the mean birth weight was 1555.41\u0026thinsp;\u0026plusmn;\u0026thinsp;576.58 g, and the mean NICU stay was 24.30\u0026thinsp;\u0026plusmn;\u0026thinsp;19.18 days. The majority of infants were male (56.2%), singleton (63.8%), and resided in urban areas (59.9%). ROP stage at the first visit ranged from 0 (38.3%) to 3 (18.0%), with 69.8% of infants showing no plus disease. Univariate analyses revealed significant differences between the treatment and non-treatment groups. Infants requiring treatment had significantly lower gestational age (28.75\u0026thinsp;\u0026plusmn;\u0026thinsp;2.12 vs 33.40\u0026thinsp;\u0026plusmn;\u0026thinsp;2.43 weeks; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, Cohen\u0026rsquo;s r\u0026thinsp;=\u0026thinsp;0.853), lower birth weight (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, large effect size), longer NICU stays (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), and higher ROP stage and plus disease prevalence (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001, Cram\u0026eacute;r\u0026rsquo;s V\u0026thinsp;\u0026gt;\u0026thinsp;0.80). Among the imaging features, PCA1 demonstrated the strongest discriminatory power (r\u0026thinsp;=\u0026thinsp;0.964, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Demographic variables showed no significant association with treatment requirement.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Multimodal Model Performance\u003c/h2\u003e \u003cp\u003eFourteen machine learning models were evaluated across five predefined multimodal feature sets incorporating PCA embeddings and/or clinical variables. Five-fold stratified cross-validation was performed on the training set (80% of data) to ensure robust performance estimation. Tree-based models (Extra Trees, Random Forest, XGBoost, LightGBM) consistently outperformed other algorithms across feature sets containing PCA embeddings (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For instance, Extra Trees and Random Forest on Set 6 (full multimodal) and Set 4 (clinical\u0026thinsp;+\u0026thinsp;PCA1\u0026ndash;3) achieved test accuracy of 0.987, AUC-ROC up to 0.999, and F1-score of 0.988, with near-perfect specificity. Cross-validation confirmed the stability of these results, with mean CV accuracy ranging from 0.984 to 0.986\u0026thinsp;\u0026plusmn;\u0026thinsp;0.007\u0026ndash;0.010 in these sets. In contrast, clinical-only sets (Set 3 and Set 5) exhibited substantially lower performance (maximum test accuracy of 0.896 and 0.948, respectively; CV accuracy 0.888\u0026ndash;0.942\u0026thinsp;\u0026plusmn;\u0026thinsp;0.018\u0026ndash;0.026), underscoring the critical contribution of vascular morphology features derived from retinal imaging.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of top-performing models on multimodal feature sets (test set \u0026amp; 5-fold CV metrics)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeature Set\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAccuracy (mean\u0026thinsp;\u0026plusmn;\u0026thinsp;std)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC-ROC (mean\u0026thinsp;\u0026plusmn;\u0026thinsp;std)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTest Accuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eTest AUC-ROC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eTest F₁-score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eTest Sensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003eTest Specificity\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSet 6\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eFull multimodal (Clinical/demographic\u0026thinsp;+\u0026thinsp;ROP\u0026thinsp;+\u0026thinsp;PCA1\u0026ndash;3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Extra Trees\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.986\u0026thinsp;\u0026plusmn;\u0026thinsp;0.007\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.995\u0026thinsp;\u0026plusmn;\u0026thinsp;0.004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.999\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2. Random Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.985\u0026thinsp;\u0026plusmn;\u0026thinsp;0.008\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.994\u0026thinsp;\u0026plusmn;\u0026thinsp;0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.998\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSet 4\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eClinical/demographic\u0026thinsp;+\u0026thinsp;PCA1\u0026ndash;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Extra Trees\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.984\u0026thinsp;\u0026plusmn;\u0026thinsp;0.009\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.994\u0026thinsp;\u0026plusmn;\u0026thinsp;0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.998\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.972\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2. Random Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.983\u0026thinsp;\u0026plusmn;\u0026thinsp;0.010\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.993\u0026thinsp;\u0026plusmn;\u0026thinsp;0.006\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.997\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSet 2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ePCA1\u0026ndash;3\u0026thinsp;+\u0026thinsp;ROP stage/zone/plus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Random Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.985\u0026thinsp;\u0026plusmn;\u0026thinsp;0.008\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.992\u0026thinsp;\u0026plusmn;\u0026thinsp;0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.990\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2. Extra Trees\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.986\u0026thinsp;\u0026plusmn;\u0026thinsp;0.007\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.993\u0026thinsp;\u0026plusmn;\u0026thinsp;0.004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.987\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.992\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSet 5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eClinical/demographic\u0026thinsp;+\u0026thinsp;ROP stage/zone/plus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. XGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.942\u0026thinsp;\u0026plusmn;\u0026thinsp;0.018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.970\u0026thinsp;\u0026plusmn;\u0026thinsp;0.012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.948\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.972\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.952\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.917\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2. Random Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.938\u0026thinsp;\u0026plusmn;\u0026thinsp;0.020\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.968\u0026thinsp;\u0026plusmn;\u0026thinsp;0.013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.935\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.968\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.942\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.889\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eSet 3\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eClinical/demographic only\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Random Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.890\u0026thinsp;\u0026plusmn;\u0026thinsp;0.025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.935\u0026thinsp;\u0026plusmn;\u0026thinsp;0.018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.896\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.938\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.927\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.861\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2. Extra Trees\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e0.888\u0026thinsp;\u0026plusmn;\u0026thinsp;0.026\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e0.933\u0026thinsp;\u0026plusmn;\u0026thinsp;0.019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.896\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.935\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.927\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.861\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Model Interpretability (SHAP Analysis, Feature importance)\u003c/h2\u003e \u003cp\u003eSHAP analysis was performed on the top-performing models (Extra Trees, Random Forest, XGBoost, and LightGBM) across feature Sets 2, 4, 5, and 6. The results consistently ranked PCA1 as the most influential feature (mean absolute SHAP value\u0026thinsp;\u0026gt;\u0026thinsp;10\u0026ndash;12), followed by ROP stage at the first visit and gestational age. Higher values of PCA1 were strongly associated with an increased likelihood of requiring treatment, reflecting more severe vascular abnormalities in the retina. Synergistic interactions were observed between PCA1 and key clinical risk factors, such as low gestational age, indicating that the models effectively captured the combined effects of imaging-derived vascular morphology and established ROP risk factors. These findings confirm that the models\u0026rsquo; predictions are grounded in clinically meaningful features rather than spurious correlations.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe results of this study demonstrate that our explainable multimodal deep learning model achieves high predictive performance for treatment-requiring retinopathy of prematurity, with Extra Trees and Random Forest classifiers yielding a test accuracy of 0.987, AUC-ROC of 0.999, and F1-score of 0.988 on the full multimodal feature set. These outcomes substantially outperform clinical-only models (maximum test accuracy 0.948) and align with or exceed recent AI advancements in ROP prediction, underscoring the value of integrating retinal imaging features\u0026mdash;particularly PCA1, which emerged as the most influential predictor in SHAP and Grad-CAM analyses (r\u0026thinsp;=\u0026thinsp;0.881 with treatment requirement).\u003c/p\u003e \u003cp\u003eOur findings are consistent with prior image-only convolutional neural network (CNN) models, which have reported accuracies exceeding 90% for detecting plus disease or treatment-requiring cases (Tan et al. [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]; Huang et al. [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]; Yenice et al. [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]). For instance, Tan et al. achieved 96.6% sensitivity in plus disease detection using DL on fundal images, while Huang et al. reported up to 96.14% sensitivity for detecting ROP (no ROP vs. stages 1\u0026ndash;2). However, multimodal approaches combining imaging with clinical data have consistently shown superior performance, particularly in predicting disease progression or reactivation post-treatment (Wu et al. [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]; He et al. [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]; Engin et al. [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]). Wu et al. developed a fusion model with an AUC of 0.822 for ROP reactivation after anti-VEGF, integrating clinical and imaging data, while He et al. achieved an AUC of 0.853 using DNN in the e-ROP study. Our model extends this paradigm by incorporating semi-supervised vessel segmentation via a U-Net-based adversarial domain-adaptation approach, addressing challenges in heterogeneous neonatal fundus images and achieving robust domain-invariant vascular feature extraction\u0026mdash;advantages not fully realized in prior studies like Redd et al. [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e], which reported an AUC of 0.960 for detecting type 1 ROP but relied on basic U-Net vessel segmentation without domain adaptation.\u003c/p\u003e \u003cp\u003eThe attention mechanism in our multi-view fusion further enhances interpretability by adaptively weighting informative retinal perspectives, a critical factor for clinical adoption. The dominance of PCA1 in SHAP analyses highlights the pivotal role of vascular morphology (tortuosity, dilation, and abnormal branching) in identifying severe ROP, consistent with findings from Krishnan et al. [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e], who identified gestational age and device type as key predictors in ML models, and Karkhaneh et al. [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e], who reported 85% sensitivity in digital imaging for referral-warranted ROP in an Iranian cohort. Grad-CAM heatmaps confirmed model attention on the posterior pole, aligning with the International Classification of ROP criteria for plus disease\u0026mdash;the strongest predictor of treatment need. This biological alignment addresses a key limitation of many DL models: reliance on non-biological artifacts, as noted in systematic reviews (Jafarizadeh et al.[\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]).\u003c/p\u003e \u003cp\u003eCompared to prior multimodal ROP models, our approach offers several advantages. While Takeda et al. [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e] achieved AUCs of 0.93\u0026ndash;0.94 using non-imaging ML for ROP occurrence prediction, and Shoeibi et al.[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e] reported 96% accuracy in XGBoost for treatment-needed ROP, these studies often lack external validation, comprehensive clinical integration, or focus on Iranian-specific cohorts (e.g., Karkhaneh et al.). By contrast, our model incorporates a broad range of demographic and neonatal variables (e.g., gestational age, NICU duration) alongside vascular features, achieving near-perfect specificity (up to 1.000) and minimizing false positives\u0026mdash;an essential consideration for reducing unnecessary examinations in vulnerable neonates, as emphasized by Engin et al. The use of 5-fold stratified cross-validation on the training set further ensures robustness and generalizability within our cohort, surpassing the internal validation in Wu et al.\u003c/p\u003e \u003cp\u003eDespite these strengths, several limitations must be acknowledged. First, the study was conducted at a single tertiary center in northeastern Iran, potentially limiting generalizability to other populations or settings with different ROP prevalence and screening practices, as highlighted in reviews by Jafarizadeh et al. External validation on multicenter or international cohorts\u0026mdash;such as those in the e-ROP study (He et al.) or global datasets (Krishnan et al.)\u0026mdash;is essential to confirm performance across diverse ethnicities and healthcare systems. Second, while the dataset was rigorously quality-controlled, neonatal fundus imaging remains technically challenging due to motion artifacts and variable illumination; future work could explore additional preprocessing techniques or larger datasets to further enhance robustness. Third, although SHAP and Grad-CAM provided strong interpretability, prospective clinical studies are needed to evaluate the model's real-world impact on screening efficiency and blindness prevention. The clinical implications of this work are significant. Current ROP screening guidelines rely primarily on gestational age and birth weight, resulting in frequent examinations of infants who ultimately do not require treatment. By achieving high accuracy and near-perfect specificity, our model could enable risk stratification, allowing clinicians to prioritize high-risk cases for timely intervention while reducing unnecessary retinal examinations\u0026mdash;echoing benefits seen in non-imaging models by Takeda et al. This approach has the potential to decrease infant stress, alleviate parental burden, and optimize resource allocation in neonatal intensive care units, particularly in resource-limited settings like Iran.In \u003cb\u003econclusion\u003c/b\u003e, this explainable multimodal deep learning model represents a promising tool for early prediction of treatment-requiring ROP. By integrating vascular morphology from retinal images with clinical data, the framework not only achieves superior predictive performance but also provides interpretable insights aligned with clinical reasoning. Future multicenter validation and prospective implementation studies, building on global AI advancements (Jafarizadeh et al.), are warranted to translate these findings into routine clinical practice, ultimately contributing to improved outcomes for preterm infants at risk of vision loss.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cul type=\"disc\"\u003e\n \u003cli\u003eAI: Artificial Intelligence\u003c/li\u003e\n \u003cli\u003eAUC-ROC: Area Under the Receiver Operating Characteristic Curve\u003c/li\u003e\n \u003cli\u003eCNN: Convolutional Neural Network\u003c/li\u003e\n \u003cli\u003eDL: Deep Learning\u003c/li\u003e\n \u003cli\u003eGrad-CAM: Gradient-weighted Class Activation Mapping\u003c/li\u003e\n \u003cli\u003eICROP: International Classification of Retinopathy of Prematurity\u003c/li\u003e\n \u003cli\u003eML: Machine Learning\u003c/li\u003e\n \u003cli\u003eNICU: Neonatal Intensive Care Unit\u003c/li\u003e\n \u003cli\u003ePCA: Principal Component Analysis\u003c/li\u003e\n \u003cli\u003eQC: Quality Control\u003c/li\u003e\n \u003cli\u003eResNet: Residual Network\u003c/li\u003e\n \u003cli\u003eRetCam: RetCam Imaging System\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eRF: Random Forest\u003c/li\u003e\n \u003cli\u003eROP: Retinopathy of Prematurity\u003c/li\u003e\n \u003cli\u003eSHAP: SHapley Additive exPlanations\u003c/li\u003e\n \u003cli\u003eSVM: Support Vector Machine\u003c/li\u003e\n \u003cli\u003eU-Net: U-shaped Network\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eVEGF: Vascular Endothelial Growth Factor\u003c/li\u003e\n \u003cli\u003eXGBoost: eXtreme Gradient Boosting\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003e This study was approved by the Mashhad University of Medical Sciences Ethics Committee (approval number: IR.MUMS.MEDICAL.REC.1403.338) and adhered to the tenets of the Declaration of Helsinki. Written informed consent was waived due to the retrospective nature of the study and the use of anonymized data.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis study received no specific funding from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors contributed equally to the study, including conceptualization, data collection, analysis, interpretation, and manuscript preparation. FS, FBA, NS, SE, and AAT each participated in all aspects of the research and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eNot applicable.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request. The complete analysis pipeline code is also available upon request from the corresponding author.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKong L, Fry M, Al-Samarraie M, Gilbert C, Steinkuller PG. An update on progress and the changing epidemiology of causes of childhood blindness worldwide. J aapos. 2012;16(6):501\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu Y, Dang Y. Recent Advances in Understanding the Role of PEST Sequence-Containing Proteins in Retinal Neovascularization. Curr Drug Targets. 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim SJ, Port AD, Swan R, Campbell JP, Chan RVP, Chiang MF. Retinopathy of prematurity: a review of risk factors and their clinical significance. Surv Ophthalmol. 2018;63(5):618\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaur K, Mikes BA. Retinopathy of Prematurity. StatPearls. Treasure Island (FL): StatPearls Publishing Copyright \u0026copy; 2025. StatPearls Publishing LLC.; 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarkhaneh R, Shokravi N. Assessment of retinopathy of prematurity among 150 premature neonates in Farabi Eye Hospital. Acta Medica Iranica. 2001:35\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNakshab M, Bayani GH, \u0026Euml;shaghi M. Prevalence of retinopathy in premature neonates in neonatal intensive care unit of Boali sina hospital in 2001. J Mazandaran Univ Med Sci. 2003;13(39):63\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarkhaneh R, Riazi Esfahani M, Ghojehzadeh L, Kadivar M, Nayeri F, Chams H. Nili Ahmadabadi M. Incidence and risk factors of retinopathy of prematurity. Bina J Ophthalmol. 2005;11(1):81\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMansouri MR, Kadivar M, Karkhaneh R, Riazi Esfahani M, Nili Ahmadabadi M, Faghihi H, Mirshahi A, Sadat-Nayeri F, Farahvash MS, Tabatabaei A, Adelpour A. Prevalence and risk factors of retinopathy of prematurity in very low birth weight or low gestational age infants. Bina J Ophthalmol. 2007;12(4):428\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaderian G, Iranpour R, Mohammadizadeh M, Fazel Najafabadi F, Badiei Z, Naseri F, et al. The frequency of retinopathy of prematurity in premature infants referred to an ophthalmology clinic in Isfahan. J Isfahan Med Sch. 2011;29(128):126\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFayyazi A, Heidarzadeh M, Fayzalahzadeh H, Golzar A, Sadegi K. Prevalence of retinopathy of prematurity in preterm infant hospitalized in Tabriz Alzahra Hospital's NICU. Med J tabriz Univ Med Sci. 2009;30(4):63\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSadeghi K, Heidary E, Hashemi F, Heydarzadeh M. Incidence and risk factors of retinopathy of prematurity. Med J Tabriz Univ Med Sci. 2009;30(2):73\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRiazi-Esfahani M, Alizadeh Y, Karkhaneh R, Mansouri MR, Kadivar M, Ahmadabadi MN, Nayeri F. Retinopathy of prematurity: single versus multiple-birth pregnancies. J ophthalmic Vis Res. 2008;3(1):47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarkhaneh R, Mousavi SZ, Riazi-Esfahani M, Ebrahimzadeh SA, Roohipoor R, Kadivar M, Ghalichi L, Mohammadi SF, Mansouri MR. Incidence and risk factors of retinopathy of prematurity in a tertiary eye hospital in Tehran. Br J Ophthalmol. 2008;92(11):1446\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhalichi L. Incidence, severity and risk factors for retinopathy of prematurity in premature infants with late retinal examination. Bina J Ophthalmol. 2008 Jul 10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhatami SF, YOUSEFI AE, FATEHI BG, MAMOURI GA. Retinopathy of prematurity among 1000\u0026ndash;2000 gram birth weight newborn infants.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaderian G, Parvaresh M, Rismanchiyan A, Sajadi V. Refractive errors after laser therapy for retinopathy of prematurity. Bina J Ophthalmol. 2009;15(1):13\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFouladinejad M, Motahari MM, Gharib MH, Sheishari F, Soltani MO. The prevalence, intensity and some risk factors of retinopathy of premature newborns in Taleghani Hospital, Gorgan, Iran.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaeidi R, Hashemzadeh A, Ahmadi S. RAHMANI S. Prevalence and predisposing factors of retinopathy of prematurity in very low-birth-weight infants discharged from NICU.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMousavi SZ, Karkhaneh R, Riazi-Esfahani M, Mansouri MR, Roohipoor R, Ghalichi L, Kadivar M, Nili-Ahmadabadi M, Naieri F. Retinopathy of prematurity in infants with late retinal examination. J ophthalmic Vis Res. 2009;4(1):24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNADERIAN GA, MOULAVI VH, Hadipour M, Sajadi V. Prevalence and risk factors for retinopathy of prematurity in Isfahan.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBayat-Mokhtari M, Pishva N, Attarzadeh A, Hosseini H, Pourarian S. Incidence and risk factors of retinopathy of prematurity among preterm infants in Shiraz/Iran. Iran J Pediatr. 2010;20(3):303.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEbrahim M, Ahmad RS, Mohammad M. Incidence and risk factors of retinopathy of prematurity in Babol, North of Iran. Ophthalmic Epidemiol. 2010;17(3):166\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMOUSAVI SZ, RIAZI EM, Rouhipour R, JABARVAND M, Ghalichi L, NILI AM, GHASEMI F, AALAMI HZ, Karkhaneh R. Characteristics of advanced stages of retinopathy of prematurity.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhaseminejad A, Niknafs P. Distribution of retinopathy of prematurity and its risk factors. Iran J Pediatr. 2011;21(2):209.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAfarid M, Hosseini H, Abtahi B. Screening for retinopathy of prematurity in South of Iran. Middle East Afr J Ophthalmol. 2012;19(3):277\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFeghhi M, Altayeb SM, Haghi F, Kasiri A, Farahi F, Dehdashtyan M, Movasaghi M, Rahim F. Incidence of retinopathy of prematurity and risk factors in the south-western region of Iran. Middle East Afr J Ophthalmol. 2012;19(1):101\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbrishami M, Maemori GA, Boskabadi H, Yaeghobi Z, Mafi-Nejad S, Abrishami M. Incidence and risk factors of retinopathy of prematurity in mashhad, northeast iran. Iran Red Crescent Med J. 2013;15(3):229.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKazem Sabzehei M, Afje h SA, Farahani AD, Shamshiri AR, Esmaili F. Retinopathy of Prematurity: Incidence, Risk Factors, and Outcome. Archives Iran Med (AIM). 2013;16(9).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmadpour-kacho M, Pasha YZ, Rasoulinejad SA, Hajiahmadi M, Pourdad P. Correlation between retinopathy of prematurity and clinical risk index for babies score. Tehran Univ Med J. 2014;72(6).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRasoulinejad SA, Montazeri M. Retinopathy of prematurity in neonates and its risk factors: a seven year study in northern Iran. Open Ophthalmol J. 2016;10:17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBroxterman EC, Hug DA. Retinopathy of Prematurity: A Review of Current Screening Guidelines and Treatment Options. Mo Med. 2016;113(3):187\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaur K, Gurnani B, Kannusamy V, Yadalla D. A tale of orbital cellulitis and retinopathy of prematurity in an infant: First case report. Eur J Ophthalmol. 2022;32(6):Np20\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHutchinson AK, Melia M, Yang MB, VanderVeen DK, Wilson LB, Lambert SR. Clinical Models and Algorithms for the Prediction of Retinopathy of Prematurity: A Report by the American Academy of Ophthalmology. Ophthalmology. 2016;123(4):804\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalmer EA, Flynn JT, Hardy RJ, Phelps DL, Phillips CL, Schaffer DB, et al. Incidence and early course of retinopathy of prematurity. The Cryotherapy for Retinopathy of Prematurity Cooperative Group. Ophthalmology. 1991;98(11):1628\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRevised indications for the. treatment of retinopathy of prematurity: results of the early treatment for retinopathy of prematurity randomized trial. Arch Ophthalmol. 2003;121(12):1684\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOphthalmology So P, AAo O, AAo. Ophthalmology AAfP, Strabismus. Screening Examination of Premature Infants for Retinopathy of Prematurity. Pediatrics. 2006;117(2):572\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDesai S, Athikarisamy SE, Lundgren P, Simmer K, Lam GC. Validation of WINROP (online prediction model) to identify severe retinopathy of prematurity (ROP) in an Australian preterm population: a retrospective study. Eye (Lond). 2021;35(5):1334\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShoeibi N, Ameri N, Hoseinkhani MR, Ansari Astaneh MR, Motamed Shariati M, Hosseini SM, Abrishami M, Abrishami M, Zamani G, Heidarzadeh HR. Deep learning algorithms for timely diagnosis of retinopathy of prematurity requiring treatment. Eye 2025 Nov 3:1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J, Ji J, Zhang M, Lin JW, Zhang G, Gong W, Cen LP, Lu Y, Huang X, Huang D, Li T. Automated explainable multidimensional deep learning platform of retinal images for retinopathy of prematurity screening. JAMA Netw open. 2021;4(5):e218758.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu R, Zhang Y, Huang P, Xie Y, Wang J, Wang S, Lin Q, Bai Y, Feng S, Cai N, Lu X. Prediction of Reactivation After Antivascular Endothelial Growth Factor Monotherapy for Retinopathy of Prematurity: Multimodal Machine Learning Model Study. J Med Internet Res. 2025;27:e60367.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJafarizadeh A, Maleki SF, Pouya P, Sobhi N, Abdollahi M, Pedrammehr S, Lim CP, Asadi H, Alizadehsani R, Tan RS, Islam SM. Current and future roles of artificial intelligence in retinopathy of prematurity. Artif Intell Rev. 2025;58(6):1\u0026ndash;55.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y, Yang L, Chen J, Fredericksen M, Hughes DP, Chen DZ. Deep adversarial networks for biomedical image segmentation utilizing unannotated images. InInternational conference on medical image computing and computer-assisted intervention. 2017 Sep 4 (pp. 408\u0026ndash;416). Cham: Springer International Publishing.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTαν Z, Si\u0026micro;kiν S, Lαi C, Dαi S. Deeπ leαρνiνg αlgoρiτh\u0026micro; φoρ αuτo\u0026micro;ατeδ δiαgνoσiσ oφ ρeτiνoπατhy oφ πρe\u0026micro;ατuρiτy πluσ δiσeασe. Tρανσlατioναl viσioν σcieνce τechνology. 2019;8(6):23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang YP, Basanta H, Kang EY, Chen KJ, Hwang YS, Lai CC, Campbell JP, Chiang MF, Chan RV, Kusaka S, Fukushima Y. Automated detection of early-stage ROP using a deep convolutional neural network. Br J Ophthalmol. 2021;105(8):1099\u0026ndash;103.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKıran Yenice E, Kara C, Erdaş \u0026Ccedil;B. Automated detection of type 1 ROP, type 2 ROP and A-ROP based on deep learning. Eye. 2024;38(13):2644\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe D, Luo X, Ying B, Quinn GE, Baumritter A, Chen Y, Ying GS, He L. Machine learning models for predicting treatment-requiring retinopathy of prematurity in the e-ROP study. Translational Vis Sci Technol. 2025;14(8):14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDurmaz Engin C, Ozturk T, Ozkan O, Oztas A, Selver MA, Tuzun F. Prediction of retinopathy of prematurity development and treatment need with machine learning models. BMC Ophthalmol. 2025;25(1):1\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRedd TK, Campbell JP, Brown JM, Kim SJ, Ostmo S, Chan RVP, et al. Evaluation of a deep learning image assessment system for detecting severe retinopathy of prematurity. Br J Ophthalmol. 2019;103(5):580\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKρiσhναν A, Sαl\u0026micro;αν Sαρkαρ M, Bαδαvατh L. Mαchiνe Leαρνiνg Aππlicατioνσ iν Reτiνoπατhy oφ Pρe\u0026micro;ατuρiτy Diαgνoσiσ Uσiνg τhe ROP Reτiναl I\u0026micro;αge Dατασeτ. J Neoναταl Suρg [Iντeρνeτ]. 2025 Φeβ. 7 [ciτeδ 2025 Dec. 28];14(1S):820-7. Avαilαβle φρo\u0026micro;: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehττπσ://jνeoναταlσuρg.co\u0026micro;/iνδex.πhπ/jνσ/αρτicle/vieω/1607\u003c/span\u003e\u003cspan address=\"http://hττπσ://jνeoναταlσuρg.co\u0026micro;/iνδex.πhπ/jνσ/αρτicle/vieω/1607\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKαρkhανeh R, Ah\u0026micro;αδραji A, Eσφαhανi MR, Roohiπouρ R, Dαστjανi AΦ, I\u0026micro;ανi M, Khoδαβανδe A, Eβραhi\u0026micro;iαδiβ N, Ah\u0026micro;αδαβαδi MN. The αccuραcy oφ δigiταl i\u0026micro;αgiνg iν δiαgνoσiσ oφ ρeτiνoπατhy oφ πρe\u0026micro;ατuρiτy iν Iραν: α πiloτ στuδy. Jouρναl oφ Oπhτhαl\u0026micro;ic \u0026amp; Viσioν Reσeαρch. 2019 Jαν;14(1):38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakeda Y, Kaneko Y, Sugimoto M, Yamashita H, Sasaki A, Mitsui T. Prediction Models for Retinopathy of Prematurity Using Nonimaging Machine Learning Approaches: A Regional Multicenter Study. Ophthalmol Sci. 2025;5(4):100715.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShoeibi N, Abrishami M, Hosseini SM, Ansari-Astaneh MR, Farrahi R, Gharib B, et al. Development and validation of machine learning classifiers for predicting treatment-needed retinopathy of prematurity. BMC Med Inf Decis Mak. 2025;25(1):221.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Retinopathy of Prematurity, multimodal analysis, machine learning, deep learning, fundus images","lastPublishedDoi":"10.21203/rs.3.rs-8838478/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8838478/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Retinopathy of prematurity (ROP) is a leading preventable cause of childhood blindness. Current screening guidelines, based primarily on gestational age and birth weight, result in numerous unnecessary examinations. We aimed to develop an explainable multimodal deep learning model for early prediction of treatment-requiring ROP.\u003c/p\u003e\n\u003cp\u003e‏\u003cstrong\u003eMethods:\u003c/strong\u003e In a retrospective cohort of 384 preterm infants (203 treated, 181 untreated) from a tertiary center in Iran (2021-2024), we integrated four directional fundus images, semi-supervised vessel segmentation maps (using a U-Net-based adversarial domain-adaptation approach), and comprehensive clinical/demographic data. A multi-view fusion model with an attention mechanism extracted vessel-aware features, reduced via PCA, and combined with clinical variables. Six multimodal feature sets were evaluated using 14 machine learning classifiers with 5-fold stratified cross-validation.\u003c/p\u003e\n\u003cp\u003e‏\u003cstrong\u003eResults:\u003c/strong\u003e The best-performing models (Extra Trees and Random Forest) on the full multimodal feature set achieved a test accuracy of 0.987, an AUC-ROC of 0.999, an F1-score of 0.988, and near-perfect specificity (up to 1.000). Interpretability analyses (SHAP and Grad-CAM) confirmed that predictions were primarily driven by vascular morphology features (PCA1) and posterior pole abnormalities consistent with plus disease.\u003c/p\u003e\n\u003cp\u003e‏\u003cstrong\u003eConclusions:\u003c/strong\u003e The proposed explainable multimodal model significantly outperforms clinical-only approaches and represents a promising tool for risk stratification in ROP screening. It has the potential to reduce unnecessary examinations, infant stress, and healthcare burden while facilitating timely intervention. External multicenter validation is warranted.\u003c/p\u003e","manuscriptTitle":"Explainable Multimodal Deep Learning Model for Early Prediction of Treatment-Requiring Retinopathy of Prematurity","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-13 08:57:13","doi":"10.21203/rs.3.rs-8838478/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-26T07:07:50+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-25T13:00:07+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-16T12:30:31+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-16T03:23:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"137377320909118470925249392805804376123","date":"2026-03-16T02:44:14+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-13T08:43:43+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"67661227733874064822987059868053409429","date":"2026-03-12T09:16:43+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"65654562001512699172287717813088291394","date":"2026-03-11T10:58:58+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"60837422618625309102854515519747550343","date":"2026-03-09T10:57:21+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-09T06:23:38+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-12T11:33:46+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-11T08:27:26+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-11T08:21:53+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Informatics and Decision Making","date":"2026-02-10T08:06:06+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8ef8fd61-7673-4539-8f3c-1c8f62b36751","owner":[],"postedDate":"March 13th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-07T05:38:18+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-13 08:57:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8838478","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8838478","identity":"rs-8838478","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00