Predicting Malignancy in Breast Lesions: Enhancing Accuracy with Fine-Tuned Convolutional Neural Network Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting Malignancy in Breast Lesions: Enhancing Accuracy with Fine-Tuned Convolutional Neural Network Models Li Li, Changjie Pan, Ming Zhang, Dong Shen, Guangyuan He, Mingzhu Meng This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3937557/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 11 Nov, 2024 Read the published version in BMC Medical Imaging → Version 1 posted 10 You are reading this latest preprint version Abstract Objectives. This study aimed to explore which convolutional neural network (CNN) model is best for predicting the likelihood of malignancy on dynamic contrast-enhanced breast magnetic resonance imaging (DCE-BMRI). Materials and Methods. A total of 273 benign (benign group) and 274 malignant lesions (malignant group) were obtained, and randomly divided into a training set (benign group: 246 lesions, malignant group: 245 lesions) and a testing set (benign group: 28 lesions, malignant group: 28 lesions) in a 9:1 ratio. An additional 53 lesions from 53 patients were designated as the validation set. Five models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) were evaluated. The metrics for model performance evaluation included accuracy (Ac) in the training and testing sets, and precision (Pr), recall rate (Rc), F1 score (F1), and area under the receiver operating characteristic curve (AUC) in the validation set. Results. Accuracies of 1.0 were achieved on the training set by all five fine-tuned models (S1-5), with model S4 demonstrating the highest test accuracy at 0.97. Additionally, S4 showed the lowest loss value in the testing set. The S4 model also attained the highest AUC (Area Under the Curve) of 0.89 in the validation set, marking a 13% improvement over the VGG19 model. Notably, the AUC of S4 for BI-RADS 3 was 0.90 and for BI-RADS 4 was 0.86, both significantly higher than the 0.65 AUC for BI-RADS 5. Conclusion. The S4 model we propose emerged as the superior model for predicting the likelihood of malignancy in DCE-BMRI and holds potential for clinical application in patients with breast diseases. However, further validation is necessary, underscoring the need for additional data. BI-RADS Convolutional Neural Networks Deep transfer learning Breast lesions Magnetic resonance imaging Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 1 Introduction In 2020, breast cancer accounted for 2.3 million new cases among women, surpassing lung cancer as the most commonly diagnosed cancer in this demographic, with a prevalence of 11.7%[ 1 ].Despite a continued decline in breast cancer mortality in the United States, which saw a 40% reduction between 1989 and 2017, there was a notable 0.3% annual increase in incidence rates over a five-year period (2012–2016). This rise is primarily attributed to increasing rates of local stage and hormone receptor-positive diseases[ 2 ]. While China’s cancer incidence rate remains lower than those of the United Kingdom and the United States, the expected rise in cancer cases in the coming years is a concern. This anticipated increase is due to an aging population, population growth, and more prevalent westernized lifestyles[ 3 ]. Consequently, there has been a surge in female breast cancer cases in China, mirroring a trend also observed in developed countries like the USA and UK [ 4 ]. Recently, the diagnosis of breast cancer has become increasingly complex, owing to a more comprehensive understanding of the hallmark characteristics of breast tumors[ 5 , 6 ]. Early detection and accurate diagnosis are imperative for effective treatment and improved outcomes, ultimately contributing to a reduction in the breast cancer mortality rate. The Breast Imaging Reporting and Data System (BI-RADS) prototype, first published by the American College of Radiology (ACR) in 1993, addressed the lack of uniformity in mammography reporting and has since undergone several revisions[ 7 – 9 ]. BI-RADS has gained wide acceptance among clinicians and radiologists. The fifth edition, updated in 2013, further clarified the professional terminology for breast cancer and developed standardized imaging reports. Efforts were made to ensure compatibility across all three imaging lexicons ((mammography, ultrasound, and magnetic resonance imaging (MRI)) by uniformly referring to lesions. BI-RADS category 3 (probably benign) findings warrant short-term follow-up, while category 4 and 5 lesions require histopathological examination for definitive conclusions. Most BI-RADS 3, 4, and 5 lesions are diagnosed as benign. Besides the additional costs incurred from further examinations and procedures, patients also experience increased psychological stress, anxiety, and inconvenience. Consequently, reclassifying more benign lesions as BI-RADS category 1 or 2 could create a win-win situation for both patients and health systems. Dynamic contrast-enhanced breast magnetic resonance imaging (DCE-BMRI) may offer better differentiation between malignant and benign features, potentially reducing unnecessary imaging follow-ups and benign biopsies, although current literature lacks consistency in this regard. Presently, various methods are employed for early detection of breast cancer. While many screening methods are in use, they each have limitations. In this study, we utilized a state-of-the-art convolutional neural network (CNN) to classify breast lesions, assessing its potential to improve diagnostic accuracy. Deep transfer learning (DTL) techniques have been successfully applied in medical image analysis[ 10 – 13 ], with pre-trained neural networks being used to differentiate between benign and malignant breast lesions in ultrasound images[ 14 ]. However, the application of these techniques to DCE-BMRI images has been limited. Therefore, the primary aim of this study was to determine which pre-trained model most effectively predicts the likelihood of malignancy in DCE-BMRI. 2 Materials and Methods 2.1 Dataset 1: Training and Testing Set We collected data from 530 patients with complete DCE-BMRI and pathological information, spanning January 2017 to December 2020. This included 17 patients with bilateral lesions (both benign and malignant lesions on one side). All lesions were pathologically confirmed and categorized into benign or malignant groups. These were then randomly assigned to a training set (benign: 246 lesions, malignant: 245 lesions) and a testing set (benign: 28 lesions, malignant: 28 lesions) in a 9:1 ratio (refer to Fig. 1 ). Variables such as age, pathological type, and tumor diameter were compared between groups. Table 1 details the pathological distribution of breast lesions. Inclusion criteria were: ① Patients not subjected to preoperative chemotherapy or chemoradiotherapy before MRI, ② Absence of puncture or surgical procedures prior to MRI. Due to space constraints, clinical presentation details are omitted. To minimize bias from bilateral lesions, only unilateral DCE-BMRI images were used. Table 1 The pathological distribution of breast lesions Pathological diagnosis Lesions Percent (%) Malignant lesions Invasive ductal carcinoma 220 80.29 Intraductal carcinoma 33 12.04 Invasive lobular carcinoma 7 2.55 Mucinous carcinoma 10 3.65 Lymphoma 1 0.36 Papillary carcinoma 3 1.09 Total 274 100.00 Benign lesions Cyst 26 9.52 Adenosis 42 15.38 Fibroadenoma 176 64.47 Chronic inflammation 6 2.20 Intraductal papilloma 20 7.33 Lobular tumor 3 1.10 Total 273 100.00 2.2 Dataset 2: Validation Set Simultaneously, 53 lesions from 53 patients were included as Dataset 2, using the same MRI scanner as Dataset 1, but unseen during training. Dataset 2 comprised three subsections: BI-RADS 3, 4, and 5 (see Fig. 1 ). Surgical pathology confirmed final diagnoses in cases where percutaneous biopsy indicated high risk. Absence of surgery with imaging stability was deemed indicative of no associated cancer. Follow-up adhered to referenced criteria[ 9 , 15 ]. Correct classification of a lesion required accurate classification in six out of ten images. Table 2 lists the specific details of Dataset 2. Table 2 Patients in dataset 2 Category N Confirmed follow up B M pathologically BI-RADS 3 10 10 4 16 BI-RADS 4 10 10 14 6 BI-RADS 5 3 10 13 0 2.3 MRI Techniques We employed two 3T MRI scanners with dedicated breast coils in a prone position. Gd-DTPA (0.1 mmol/kg, 2.50 mL/s) was injected through the elbow vein. The process involved six dynamic enhancement phases (one pre-contrast, five post-contrast). MRIs were conducted preoperatively and before initiating therapy. Detailed scanning parameters are outlined in Table 3 . Table 3 Scan parameters for the two magnetic resonance scanners Parameter Philips Achieva GE Healthcare Field strength 3.0T 3.0T No. of coil channels 8 8 Acquisition plane Axial Axial Pulse sequence 3D gradient echo (Thrive) Enhanced fast gradient echo 3D Repetition time (ms) 5.5 9.6 Echo time (ms) 2.7 2.1 Flip angle 10° 10° No. of postcontrast sequence 5 5 Fat suppression Yes Yes Scan time 570s 500s 3D, three dimensional; ms, millisecond; s: second. 2.4 Readers Five experienced radiologists from our department, each with over five years of breast MRI interpretation experience and specialized training in breast imaging, were enlisted. MRI image analyses were conducted using the GOLDPACS viewer ( www.jinpacs.com ). 2.5 Proposed model The study utilized a computer equipped with an Intel (R) Core (TM) i7-10700F, NVIDIA RTX 2060 GPU, running on Windows 10 Enterprise 64-bit with 6GB RAM. All extraneous programs were closed during model operation. Each network underwent identical data testing and training for consistent comparison. Malignant images were identified based on a threshold of ≥ 0.5, while images below this threshold were considered benign. We selected five commonly used pretrained models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) and employed five-fold cross-validation to assess model performance, selecting the best-performing model. This cross-validation process was then applied to Dataset 2. Additionally, we enhanced model performance using various fine-tuning strategies. The architecture of the proposed DTL with the five models for breast lesion classification is depicted in Fig. 2 . Initially, the images underwent random shuffling. Data augmentation techniques (rotation, shear range, zoom range, and horizontal flip) were applied prior to training. The binary cross-entropy loss function was used, and the training process was optimized using the Adam optimizer with a learning rate of 0.001. Our model required 200 epochs for training on DCE-BMRI images, with a batch size of 64 images. Activation functions included ReLU and sigmoid, as detailed in Equations 1 and 2 $$\text{R}\text{e}\text{l}\text{u}\left(\text{x}\right)=\text{f}\left(\text{x}\right)=\left\{\begin{array}{c}max(0,x), x\ge 0\\ 0, x<0\end{array}\right.$$ 1 $$\text{S}\text{i}\text{g}\text{m}\text{o}\text{i}\text{d}\left(\text{x}\right)=\text{f}\left(\text{x}\right)=\frac{1}{1+{\text{e}}^{-\text{x}}}$$ 2 2.6 Evaluation metrics We assessed the effectiveness of Deep Transfer Learning (DTL) models using five performance metrics: accuracy (Ac), precision (Pr), recall rate (Rc), F1 score (F1), and the area under the receiver operating characteristic curve (AUROC)[ 16 ]. For this analysis, cases were classified as either malignant or benign, representing positive and negative cases, respectively. True positives (TP) and true negatives (TN) denote the proportion of correctly diagnosed malignant and benign cases. False positives (FP) and false negatives (FN) indicate lesions misdiagnosed as benign and malignant, respectively. The formulas for these metrics are as follows: $$\text{A}\text{c}=\frac{\text{T}\text{P}+\text{T}\text{N}}{\text{T}\text{P}+\text{T}\text{N}+\text{F}\text{P}+\text{F}\text{N}}$$ 3 $$\text{P}\text{r}=\frac{\text{T}\text{P}}{\text{T}\text{P}+\text{F}\text{P}}$$ 4 $$\text{R}\text{c}=\frac{\text{T}\text{P}}{\text{T}\text{P}+\text{F}\text{N}}$$ 5 $$\text{F}1=\frac{2\times \text{P}\text{r}\times \text{R}\text{c}}{\text{P}\text{r}+\text{R}\text{c}}$$ 6 Notably, the accuracy metric (Ac) does not account for data distribution. The F1 score is a balanced measure that considers both precision and recall, making it particularly useful in datasets with imbalanced classes. 2.7 Statistical analysis Statistical analyses were conducted using SPSS 23.0 software (IBM). For data adhering to a normal distribution, counting data were presented as mean ± standard deviation ( \(\stackrel{-}{\text{x}}\) ± s). One-way analysis of variance (ANOVA) was employed for variance analysis between groups. The Mann-Whitney U test was applied for data not meeting the normal distribution criteria. The chi-square test was utilized for comparing frequency counts between malignant and benign groups in the datasets (training and testing sets). A P-value of < 0.05 (two-tailed) was considered statistically significant. 3 Results 3.1 Age and Lesion Diameter Age and lesion diameter did not conform to normal distribution. The age difference between the malignant group (46.40 ± 10.90 years) and the benign group (44.84 ± 10.20 years) was not statistically significant (P = 0.136). However, lesion diameters were significantly smaller in the malignant group (25.06 ± 11.54 mm) compared to the benign group (33.44 ± 16.69 mm) (P < 0.001). No significant variance was observed in lesion distribution between the training and testing sets across both groups (P = 0.988). 3.2 Cross validation We evaluated five models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) through five-fold cross-validation in Dataset 1 (see Table 4 for results). The DenseNet201 and MobileNetV2 models achieved perfect accuracy (1.00) in the training set, but their testing set accuracies were lower at 0.91 and 0.88, respectively, both below VGG19’s 0.96. Despite similar architectures, VGG19 outperformed VGG16 (0.91). However, both VGG16 and VGG19 exhibited premature loss increases with epoch advancement, indicating non-convergence on Dataset 1 and potential overfitting. Similar trends were observed for MobileNetV2 and DenseNet201. ResNet50 showed the lowest accuracy among the models (0.92 training, 0.67 testing). Figures 3 and 4 illustrate the learning curves and heat maps. Table 4 The results of the five-fold cross-validation Folds Accuracies of the training set Accuracies of the testing set model1 model2 model3 model4 model5 model1 model2 model3 model4 model5 Fold1 1.00 1.00 1.00 0.92 1.00 0.91 0.96 0.91 0.67 0.88 Fold2 1.00 1.00 1.00 0.93 0.99 0.90 0.96 0.91 0.67 0.87 Fold3 1.00 1.00 1.00 0.92 1.00 0.91 0.96 0.91 0.67 0.88 Fold4 1.00 1.00 1.00 0.92 0.99 0.90 0.96 0.91 0.67 0.87 Fold5 1.00 1.00 1.00 0.93 1.00 0.91 0.96 0.91 0.67 0.88 Note. Model1, VGG16; Model2, VGG19; Model3, DenseNet201; Model4, ResNet50; Model5, MobileNetV2. 3.3 Fine-tuning strategy Given these findings, we focused on enhancing the VGG19 model through five distinct fine-tuning strategies (Fig. 5 ). The fine-tuning involved activating neural network parameters for training, while keeping certain layers frozen. We noted that the accuracy achieved was 1.0, for all five fine-tuning models(S1-5) on the training set, but S4 obtained the highest test accuracy of 0.97 on the testing set. In addition, the loss value was the lowest in the testing set for S4. These results reveal that the S4 model has a better generalization ability than the other fine-tuned models. 3.4 ROC Analysis on Validation Set Analysis of the Receiver Operating Characteristic (ROC) curve for the five models on the validation set revealed VGG19 as the highest performer (AUC 0.92), yet the validation set AUC was only 0.76. Among the fine-tuned models, S4 attained the highest AUC (0.89) on the validation set, marking a 13% improvement over the original VGG19 (Fig. 6 ). Further analysis of S4 across BI-RADS categories 3, 4, and 5 showed notably higher AUCs for BI-RADS 3 (0.90) and 4 (0.86) compared to 5 (0.65) (Fig. 7 ). 3.5 Classification Reports on Validation Set Classification reports for the five models and S1-5 strategies are provided in Table 5 . For the validation set, VGG19 achieved higher performance metrics (Pr 0.75, Rc 0.76, F1 0.73, AUC 0.76) compared to the other models. Strategy S4 outperformed all others on the validation set with Pr 0.89, Rc 0.88, F1 0.87, and AUC 0.89. Table 5 Classification report of deep transfer learning models in validation set DTL models Pr Rc \(\text{f}1\) AUC group1 group2 avg group1 group2 avg group1 group1 avg model1 0.93 0.53 0.73 0.60 0.91 0.75 0.73 0.67 0.70 0.76 model2 0.91 0.59 0.75 0.63 0.90 0.76 0.75 0.71 0.73 0.73 model3 0.91 0.52 0.71 0.59 0.88 0.74 0.72 0.65 0.69 0.71 model4 0.90 0.39 0.65 0.53 0.84 0.68 0.67 0.53 0.60 0.65 model5 0.91 0.55 0.73 0.61 0.89 0.75 0.73 0.68 0.71 0.73 S1 0.88 0.63 0.75 0.65 0.87 0.76 0.74 0.73 0.74 0.75 S2 0.95 0.60 0.77 0.64 0.94 0.79 0.77 0.73 0.75 0.77 S3 0.91 0.60 0.76 0.63 0.91 0.76 0.74 0.71 0.73 0.75 S4 0.98 0.79 0.89 0.78 0.98 0.88 0.87 0.88 0.87 0.89 S5 0.94 0.59 0.77 0.64 0.93 0.79 0.76 0.73 0.75 0.77 Note. Group1, benign group; Group2, malignant group; Avg, average; Model1, VGG16; Model2, VGG19; Model3, DenseNet201; Model4, ResNet50; Model5, MobileNetV2. 4 Discussion In this study, we evaluated five pre-trained convolutional neural network models using a 5-fold cross-validation method on our DCE-BMRI dataset. Our objective was to identify the best-performing model, which we defined as the one excelling across all predefined evaluation criteria. Following this, we fine-tuned the chosen model to enhance its performance further and tested its generalization capability on a validation set. Our findings indicated that the VGG19 model demonstrated superior performance, achieving accuracies of 1.00 and 0.96 on the training and testing sets, respectively. Moreover, VGG19 achieved the highest Area Under the Curve (AUC) of 0.92 on the first validation set, but this dropped to 0.76 on a subsequent validation set, suggesting limitations in its generalization ability. Previous research supports the notion that fine-tuning can enhance the accuracy and precision of such models[ 17 – 19 ]. Consequently, we developed five distinct fine-tuning strategies for VGG19. Strategy S4 emerged as the most successful, yielding the highest test accuracy (0.97) and the lowest test loss on the validation set, indicating a superior generalization capability compared to the other strategies. When comparing the AUC scores of strategies S1-5 on the validation set, S4 again scored highest with an AUC of 0.89. These results are promising for advancing the accuracy of medical image classification diagnostics. We also delved into whether the S4 model exhibited different AUC scores across BI-RADS categories 3, 4, and 5. Interestingly, S4 performed best in BI-RADS 3 (AUC 0.90), followed by BI-RADS 4 (AUC 0.86), and showed the least performance in BI-RADS 5 (AUC 0.65). The BI-RADS 3 category, typically applied when the likelihood of cancer is less than 2%, aims to minimize unnecessary biopsies for pathologically benign findings. However, patient compliance with follow-up MRI recommendations every six months is notably low in this category[ 20 ]. The BI-RADS system is set to evolve with new breast imaging modalities. Key areas for improvement include expanding the lexicon for common findings and clarifying the application of Category 3[ 21 ]. BI-RADS 3 represents a significant portion (13.9%) of diagnostic exams [ 22 ], often leading to follow-up procedures for patients classified as ‘probably benign’, yet compliance with these follow-up recommendations remains a challenge [ 20 ]. This lack of compliance raises concerns about the clinical and economic implications, particularly regarding the resolution time and outcomes for patients in this category. In a particular study, only 1.4% of BI-RADS 3 lesions were found to be malignant, including two cases of delayed diagnosis at 13.2 and 33.2 months, respectively. The incidence of delayed diagnosis due to additional MRI-detected lesions during follow-up was notably low (0.7%), consisting exclusively of T1N0 contralateral cancers. This finding suggests that annual follow-up may suffice for BI-RADS 3 lesions identified by MRI before surgery[ 15 ]. Consequently, accurately distinguishing between benign and malignant lesions in BI-RADS 3 is crucial. Our study potentially offers significant benefits to a substantial number of patients diagnosed with BI-RADS category lesions using DCE-BMRI imaging. BI-RADS category 4 lesions are associated with a high likelihood of malignancy, with estimates ranging from 2–95%. The BI-RADS 4 classification, to a degree, is subjective; the outcomes of biopsies in this category vary significantly, and the rate of cancer detection relative to the number of biopsies performed is relatively low (17.8%)[ 15 ]. Moreover, unnecessary biopsies can result in a range of adverse effects, including pain, fear, emotional distress, and financial costs. Breast MRI is highly sensitive, yet it often presents a challenge in differentiating between atypical malignant and benign lesions, leading to potential overclassification in the BI-RADS 4 category and subsequent invasive biopsies. The wide range of positive predictive values for MRI-guided biopsies (2.5–84.0%) [ 23 , 24 ], indicates that many patients undergo unnecessary procedures. Indeed, numerous women subjected to biopsies for benign findings endure unnecessary discomfort, expenses, potential complications, cosmetic alterations, and anxiety [ 25 ]. Identifying predictors of benign BI-RADS 4 masses, therefore, could be highly beneficial. Efforts to enhance the assessment of BI-RADS 4 lesions could improve the identification of benign lesions, thereby reducing the frequency of unnecessary biopsies. Some researchers have developed predictive models based on imaging features or multiparameter MRI data to better evaluate BI-RADS 4 lesions, though these models typically rely on traditional imaging features subjectively defined by radiologists[ 26 ]. The BI-RADS 5 category is applied when imaging findings suggest a malignancy probability of 95% or higher. According to MC et al.[ 27 ], the positive predictive value of BI-RADS 5 assessments is only 71.4%, indicating that not all lesions classified as BI-RADS 5 are malignant[ 28 ], and surgery is often recommended for this category. It is recognized that a single imaging finding rarely confers such a high risk of malignancy; rather, a combination of features is necessary to elevate a lesion to Category 5[ 7 ]. However, it's important to acknowledge that even when tissue samples from molecular biopsies are used, they may not fully represent the entire lesion, as biopsies often target only a small, specific area of a heterogeneous lesion, introducing a bias in lesion selection. Rc also known as sensitivity, measures a classifier's completeness. A lower Rc value indicates the classifier's limited capability in handling large FP values. Recent publications have led to the introduction of new and updated performance benchmarks in the latest edition, replacing outdated metrics. The recall rate benchmark has been revised accordingly. Initially, about half of all radiologists were unable to meet the 10% benchmark for recall rate, prompting a revision to a more achievable target of 12%, a standard met by over 75% of radiologists[ 7 ]. Our study showed that the S4 model exhibited the highest recall rate (0.89) among all DTL models, which is notable given the relatively limited class diversity in our dataset. This study, however, is not without limitations. Firstly, the training set contained a relatively small number of images, particularly lacking in rare lesion types. Consequently, our dataset may not fully represent the broader spectrum of breast disease patients, potentially impacting the accuracy of the DTL model. Therefore, it’s essential to conduct further analysis with more extensive datasets to comprehensively evaluate the robustness of the DTL model. Secondly, our study relied solely on static Dynamic Contrast-Enhanced Breast Magnetic Resonance Imaging (DCE-BMRI) images, excluding other routine diagnostic procedures such as clinical evaluations, breast ultrasounds, and mammography. Thirdly, we limited our investigation to just five pre-trained models; future research should explore a wider range of models to determine their robustness on larger datasets. Lastly, while this paper does not delve into the various methods of fine-tuning Convolutional Neural Network (CNN) models, these topics will be the focus of our subsequent studies. 5 Conclusions The S4 model demonstrated superior accuracy in BI-RADS categories 3 and 4 compared to category 5. This finding is significant as it could lead to a reduction in the number of follow-up sessions for BI-RADS 3 and decrease the number of unnecessary biopsies for benign lesions in BI-RADS 4. However, these results necessitate further validation. Moving forward, our focus will be on exploring more robust models and the importance of augmenting our dataset with additional data. Abbreviations MRI= Magnetic resonance imaging DL= Deep learning DTL=Deep transfer learning ROC = receiver operating characteristic AUC = area under the ROC curve BI-RADS = Breast Imaging Reporting and Data System, DCE-BMRI = dynamic contrast enhanced breast MRI Declarations Ethics approval and consent to participate This study was approved by the Second Hospital of Changzhou Affiliated to Nanjing Medical University of Chinese Medicine Ethics Review Committee (Ethics Number: [2023]KY313-01). Consent for publication This study was a retrospective analysis and informed consent was waived. Availability of data and materials The datasets generated during and/or analyzed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests. Funding This study was supported by the Program of Bureau of Science and Technology Foundation of Changzhou (No. CJ20220260). Authors’ contributions Li Li and Mingzhu Meng carried out the literature search, and designed and wrote the manuscript. Mingzhu Meng and Changjie Pan conceived of the project, and participated in its design and coordination, and helped to draft the manuscript. Ming Zhang, Dong Shen and Guangyuan He are responsible for figures processing. Both authors read and approved the final manuscript. Acknowledgments The authors wish to thank Shiquan Ge for his technical assistance in operating the Python programming code. References Sung H, Ferlay J, Siegel RL,et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209-49. DeSantis CE, Ma J, Gaudet MM,et al. Breast cancer statistics, 2019. CA Cancer J Clin. 2019;69(6):438-51. Cao W, Chen HD, Yu YW,et al. Changing profiles of cancer burden worldwide and in China: a secondary analysis of the global cancer statistics 2020. Chin Med J (Engl). 2021;134(7):783-91. RM F, YN Z, SM C,et al. Current cancer situation in China: good or bad news from the 2018 Global Cancer Statistics? Cancer Commun (Lond). 2019;39(1):22. Gao Y, Heller SL. Abbreviated and Ultrafast Breast MRI in Clinical Practice. RadioGraphics. 2020;40(6):1507-27. Berdzuli N. Breast cancer: from awareness to access. BMJ. 2023;380(290. Mercado CL. BI-RADS Update. Radiol Clin North Am. 2014;52(3):481-7. Sedgwick EL, Ebuoma L, Hamame A,et al. BI-RADS update for breast cancer caregivers. Breast Cancer Research and Treatment. 2015;150(2):243-54. Pesce K, Orruma MB, Hadad C,et al. BI-RADS Terminology for Mammography Reports: What Residents Need to Know. Radiographics. 2019;39(2):319-20. Huang Z, Zhou Q, Zhu X,et al. Batch Similarity Based Triplet Loss Assembled into Light-Weighted Convolutional Neural Networks for Medical Image Classification. Sensors (Basel). 2021;21(3):764-85. Tajbakhsh N, Jeyaseelan L, Li Q,et al. Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation. Medical image analysis. 2020;63:101693. Naqa IE. The role of machine and deep learning in modern medical physics. Med Phys. 2020;47(5):e125-e6. Zhang J, Xie Y, Wu Q,et al. Medical image classification using synergic deep learning. Medical image analysis. 2019;54:10-9. Dong F, She R, Cui C,et al. One step further into the blackbox: a pilot study of how to build more confidence around an AI-based decision system of breast nodule assessment in 2D ultrasound. European radiology. 2021;31(7):4991-5000. Lee SE, Lee JH, Han K,et al. BI-RADS category 3, 4, and 5 lesions identified at preoperative breast MRI in patients with breast cancer: implications for management. Eur Radiol Exp. 2020;30(5):2773-81. Wang Z, Li X, Yao M,et al. A new detection model of microaneurysms based on improved FC-DenseNet. Sci Rep. 2022;12(1):950. Tan T, Li Z, Liu H,et al. Optimize Transfer Learning for Lung Diseases in Bronchoscopy Using a New Concept: Sequential Fine-Tuning. IEEE J Transl Eng Health Med. 2018;6:1800808. Ahamed KU, Islam M, Uddin A,et al. A deep learning approach using effective preprocessing techniques to detect COVID-19 from chest CT-scan and X-ray images. Comput Biol Med. 2021;139:105014. Montaha S, Azam S, Rafid A,et al. BreastNet18: A High Accuracy Fine-Tuned VGG16 Model Evaluated Using Ablation Study for Diagnosing Breast Cancer from Enhanced Mammography Images. Biology (Basel). 2021;10(12):1347. Ambinder EB, Myers K, Panigrahi B,et al. Breast MRI BI-RADS 3: Impact of Patient-Level Factors on Compliance With Short-Term Follow-Up. J Am Coll Radiol 2020;17(3):377-83. Eghtedari M, Chong A, Rakow-Penner R,et al. Current Status and Future of BI-RADS in Multimodality Breast Imaging, From the AJR Special Series on Radiology Reporting and Data Systems. American Journal of Roentgenology. 2020:[published online]. Elezaby M, Li G, Bhargavan-Chatfield M,et al. ACR BI-RADS assessment category 4 subdivisions in Diagnostic Mammography: Utilization and Outcomes in the National Mammography Database. Radiology. 2018;287(2):416-22. RM S, ES B, M E,et al. Utility of BI-RADS Assessment Category 4 Subdivisions for Screening Breast MRI. AJR American journal of roentgenology. 2017;208(6):1392-9. JR MdA, AB G, TP B,et al. Subcategorization of Suspicious Breast Lesions (BI-RADS Category 4) According to MRI Criteria: Role of Dynamic Contrast-Enhanced and Diffusion-Weighted Imaging. AJR American journal of roentgenology. 2015;205(1):222-31. Laws A, Crocker A, Dort J,et al. Improving Wait Times and Patient Experience Through Implementation of a Provincial Expedited Diagnostic Pathway for BI-RADS 5 Breast Lesions. Ann Surg Oncol 2019;36(10):3361-7. Hao W, Gong J, Wang S,et al. Application of MRI Radiomics-Based Machine Learning Model to Improve Contralateral BI-RADS 4 Lesion Assessment. Frontiers in Oncology. 2020;10:531476-84. MC M, C G, L H,et al. Positive Predictive Value of BI-RADS MR imaging. Radiology. 2012;264(1):51-8. Dao KA, Rives AF, Quintana LM,et al. BI-RADS 5: More than Cancer. Radiographics. 2020;40(5):1203-4. Additional Declarations No competing interests reported. Supplementary Files DataS1.doc Cite Share Download PDF Status: Published Journal Publication published 11 Nov, 2024 Read the published version in BMC Medical Imaging → Version 1 posted Editorial decision: Revision requested 19 Aug, 2024 Reviews received at journal 26 Jul, 2024 Reviews received at journal 21 Jul, 2024 Reviewers agreed at journal 19 Jul, 2024 Reviewers agreed at journal 19 Jul, 2024 Reviewers invited by journal 19 Apr, 2024 Editor invited by journal 15 Feb, 2024 Submission checks completed at journal 15 Feb, 2024 Editor assigned by journal 15 Feb, 2024 First submitted to journal 07 Feb, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3937557","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":273162305,"identity":"c118c38e-2a39-4927-8591-fce68812234c","order_by":0,"name":"Li Li","email":"","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Li","middleName":"","lastName":"Li","suffix":""},{"id":273162306,"identity":"4bc9e33a-eb08-4c0e-83f7-0967fff7e2a3","order_by":1,"name":"Changjie Pan","email":"","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Changjie","middleName":"","lastName":"Pan","suffix":""},{"id":273162307,"identity":"3ba48d42-9f6d-4a93-96a5-c582bc3452c2","order_by":2,"name":"Ming Zhang","email":"","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Ming","middleName":"","lastName":"Zhang","suffix":""},{"id":273162308,"identity":"e3bc06d2-0811-4556-aa8e-186ac97781c1","order_by":3,"name":"Dong Shen","email":"","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Dong","middleName":"","lastName":"Shen","suffix":""},{"id":273162309,"identity":"f7640fcc-58bd-4319-b9f7-56650d9a7bc0","order_by":4,"name":"Guangyuan He","email":"","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Guangyuan","middleName":"","lastName":"He","suffix":""},{"id":273162310,"identity":"444041d2-3d85-4887-b412-dc13a55c8277","order_by":5,"name":"Mingzhu Meng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9UlEQVRIiWNgGAWjYPACCTk29uYDDBJgTgJRWmyM+XmOJZCkJS1x5owcAyiHgBbd9t7Dr278Ocy44UDOtweWOYcZ+NmBen/uwK3F7My5NOscnsPMBgfObjeQ3HaYQbLnjQFj7xk8Wm7kmBnnSBxmMzjYu00CpMXgRo4BM2MbHi333wC1GBzmAaJnYC32BLXc4DF+nJOQJiHZxsMGsUWCkJYzOWbMOQdsDPh52MyAWtJ5JM48KzjYi0/L8TPGn3P+SNS3yT9+Ji25zVqOvz1544OfeLQAAZsEjMUMZPGAGAfwagAq/ABjMX7Ap24UjIJRMApGLAAAjTxRzWnQl3sAAAAASUVORK5CYII=","orcid":"","institution":"The Affiliated Changzhou No.2 People's Hospital of Nanjing Medical University","correspondingAuthor":true,"prefix":"","firstName":"Mingzhu","middleName":"","lastName":"Meng","suffix":""}],"badges":[],"createdAt":"2024-02-07 17:14:37","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3937557/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3937557/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12880-024-01484-1","type":"published","date":"2024-11-11T15:58:13+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":51387675,"identity":"81813d1b-85f2-49e2-b02f-8a762f25f3af","added_by":"auto","created_at":"2024-02-20 18:00:10","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":253671,"visible":true,"origin":"","legend":"\u003cp\u003eDataset Structure Diagram\u003c/p\u003e\n\u003cp\u003eThis figure presents a schematic representation of the dataset arrangement, illustrating how data is categorized and structured for analysis.\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/9d08d8b1bf7723dc4aa9fcc7.jpg"},{"id":51387676,"identity":"c0fc70d9-2e1e-4a4f-bf97-bb6ec77af68e","added_by":"auto","created_at":"2024-02-20 18:00:10","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":376496,"visible":true,"origin":"","legend":"\u003cp\u003eDeep Transfer Learning Network Architecture\u003c/p\u003e\n\u003cp\u003eThis figure depicts the architecture of the DTL network, highlighting its role in determining the likelihood of tumor malignancy. It emphasizes that validation sets do not have to mirror training sets and outlines the three-step data analysis process: feature extraction from the image network, training and testing of data, and data validation.\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/b75405b4d3946deb641dcf33.jpg"},{"id":51387674,"identity":"4111a5f9-b535-41d7-b652-9d3d09c92843","added_by":"auto","created_at":"2024-02-20 18:00:10","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":763110,"visible":true,"origin":"","legend":"\u003cp\u003eLearning Curves for the Five Pre-trained Models\u003c/p\u003e\n\u003cp\u003eThis figure displays learning curves for each of the five pre-trained models over various epochs, showing: a) training accuracy, b) testing accuracy, c) training loss, and d) testing loss. It notably illustrates that the VGG19 model achieved the highest accuracy in the testing set, while ResNet50 had the lowest.\u003c/p\u003e","description":"","filename":"Fig3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/a1f037525dc8ded45c7076bc.jpg"},{"id":51387677,"identity":"3359dfa1-420f-4c09-9555-f0a49eac49b6","added_by":"auto","created_at":"2024-02-20 18:00:10","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":361086,"visible":true,"origin":"","legend":"\u003cp\u003eHeatmaps of the Five Models\u003c/p\u003e\n\u003cp\u003eThe figure provides heatmaps illustrating the activated-zone boundaries for each model. It shows that the activated zones for DenseNet201 and MobileNetV2 are located outside the input image, while ResNet50's activated zone is relatively small. The heatmaps for VGG19 and VGG16 display similar locations of activation zones, with VGG19 showing greater activation.\u003c/p\u003e","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/28ce1f1212c0192b468eb5e6.jpg"},{"id":51387679,"identity":"da8b8860-acd2-4609-af97-2d6f30466898","added_by":"auto","created_at":"2024-02-20 18:00:11","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":591359,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic of Fine-Tuning Strategies for VGG19\u003c/p\u003e\n\u003cp\u003eThis figure outlines the five different fine-tuning strategies applied to the VGG19 model, detailing the number of trainable parameters, the activated layers (trainable), and the non-trainable (frozen) layers of the neural network. It also highlights the full connection (Fc) layer.\u003c/p\u003e","description":"","filename":"Figure5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/d3f723c7166ced306e09c140.jpg"},{"id":51387681,"identity":"63717711-e886-48f9-88d9-e39f5a2d5dc0","added_by":"auto","created_at":"2024-02-20 18:00:11","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":946189,"visible":true,"origin":"","legend":"\u003cp\u003eAUC Analysis of the Proposed S1-5 and VGG19 Models\u003c/p\u003e\n\u003cp\u003eThis figure showcases the Area Under the Curve (AUC) analyses for the proposed S1-5 strategies and the VGG19 model, allowing for a comparative assessment of their performance.\u003c/p\u003e","description":"","filename":"Figure6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/aef8b1d573b2c93284b56a74.jpg"},{"id":51387678,"identity":"fe0772dd-5e79-48c9-a007-991386e3a2a9","added_by":"auto","created_at":"2024-02-20 18:00:11","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":830080,"visible":true,"origin":"","legend":"\u003cp\u003eAUC Comparison in BI-RADS Subcategories for the S4 Model\u003c/p\u003e\n\u003cp\u003eThe figure compares the AUC scores of the S4 model across different BI-RADS categories (3, 4, and 5), offering insights into the model’s performance in each category.\u003c/p\u003e","description":"","filename":"Figure7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/ab4bbabe9ff3d0a38724f482.jpg"},{"id":69285112,"identity":"5985bf9a-59a8-4f27-8a8c-5566867eea22","added_by":"auto","created_at":"2024-11-18 19:23:57","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4849476,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/77d726a2-a17c-4287-b49f-6c7caf407b1c.pdf"},{"id":51387680,"identity":"0836b06e-94da-46ee-8780-90e8b81f43cf","added_by":"auto","created_at":"2024-02-20 18:00:11","extension":"doc","order_by":13,"title":"","display":"","copyAsset":false,"role":"supplement","size":74240,"visible":true,"origin":"","legend":"","description":"","filename":"DataS1.doc","url":"https://assets-eu.researchsquare.com/files/rs-3937557/v1/e678ba7e34e8273aacd9ecc4.doc"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting Malignancy in Breast Lesions: Enhancing Accuracy with Fine-Tuned Convolutional Neural Network Models","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eIn 2020, breast cancer accounted for 2.3\u0026nbsp;million new cases among women, surpassing lung cancer as the most commonly diagnosed cancer in this demographic, with a prevalence of 11.7%[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].Despite a continued decline in breast cancer mortality in the United States, which saw a 40% reduction between 1989 and 2017, there was a notable 0.3% annual increase in incidence rates over a five-year period (2012\u0026ndash;2016). This rise is primarily attributed to increasing rates of local stage and hormone receptor-positive diseases[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. While China\u0026rsquo;s cancer incidence rate remains lower than those of the United Kingdom and the United States, the expected rise in cancer cases in the coming years is a concern. This anticipated increase is due to an aging population, population growth, and more prevalent westernized lifestyles[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Consequently, there has been a surge in female breast cancer cases in China, mirroring a trend also observed in developed countries like the USA and UK [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Recently, the diagnosis of breast cancer has become increasingly complex, owing to a more comprehensive understanding of the hallmark characteristics of breast tumors[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Early detection and accurate diagnosis are imperative for effective treatment and improved outcomes, ultimately contributing to a reduction in the breast cancer mortality rate.\u003c/p\u003e \u003cp\u003eThe Breast Imaging Reporting and Data System (BI-RADS) prototype, first published by the American College of Radiology (ACR) in 1993, addressed the lack of uniformity in mammography reporting and has since undergone several revisions[\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. BI-RADS has gained wide acceptance among clinicians and radiologists.\u003c/p\u003e \u003cp\u003eThe fifth edition, updated in 2013, further clarified the professional terminology for breast cancer and developed standardized imaging reports. Efforts were made to ensure compatibility across all three imaging lexicons ((mammography, ultrasound, and magnetic resonance imaging (MRI)) by uniformly referring to lesions.\u003c/p\u003e \u003cp\u003eBI-RADS category 3 (probably benign) findings warrant short-term follow-up, while category 4 and 5 lesions require histopathological examination for definitive conclusions. Most BI-RADS 3, 4, and 5 lesions are diagnosed as benign. Besides the additional costs incurred from further examinations and procedures, patients also experience increased psychological stress, anxiety, and inconvenience.\u003c/p\u003e \u003cp\u003eConsequently, reclassifying more benign lesions as BI-RADS category 1 or 2 could create a win-win situation for both patients and health systems. Dynamic contrast-enhanced breast magnetic resonance imaging (DCE-BMRI) may offer better differentiation between malignant and benign features, potentially reducing unnecessary imaging follow-ups and benign biopsies, although current literature lacks consistency in this regard.\u003c/p\u003e \u003cp\u003ePresently, various methods are employed for early detection of breast cancer. While many screening methods are in use, they each have limitations. In this study, we utilized a state-of-the-art convolutional neural network (CNN) to classify breast lesions, assessing its potential to improve diagnostic accuracy. Deep transfer learning (DTL) techniques have been successfully applied in medical image analysis[\u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], with pre-trained neural networks being used to differentiate between benign and malignant breast lesions in ultrasound images[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. However, the application of these techniques to DCE-BMRI images has been limited. Therefore, the primary aim of this study was to determine which pre-trained model most effectively predicts the likelihood of malignancy in DCE-BMRI.\u003c/p\u003e"},{"header":"2 Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003e2.1 Dataset 1: Training and Testing Set\u003c/h2\u003e\n\u003cp\u003eWe collected data from 530 patients with complete DCE-BMRI and pathological information, spanning January 2017 to December 2020. This included 17 patients with bilateral lesions (both benign and malignant lesions on one side). All lesions were pathologically confirmed and categorized into benign or malignant groups. These were then randomly assigned to a training set (benign: 246 lesions, malignant: 245 lesions) and a testing set (benign: 28 lesions, malignant: 28 lesions) in a 9:1 ratio (refer to Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Variables such as age, pathological type, and tumor diameter were compared between groups. Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e details the pathological distribution of breast lesions. Inclusion criteria were: ① Patients not subjected to preoperative chemotherapy or chemoradiotherapy before MRI, ② Absence of puncture or surgical procedures prior to MRI. Due to space constraints, clinical presentation details are omitted. To minimize bias from bilateral lesions, only unilateral DCE-BMRI images were used.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eThe pathological distribution of breast lesions\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePathological diagnosis\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLesions\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePercent (%)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eMalignant lesions\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eInvasive ductal carcinoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e220\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e80.29\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIntraductal carcinoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e33\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e12.04\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eInvasive lobular carcinoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e2.55\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMucinous carcinoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e3.65\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLymphoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.36\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePapillary carcinoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e1.09\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTotal\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e274\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e100.00\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eBenign lesions\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCyst\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e26\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e9.52\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAdenosis\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e42\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e15.38\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFibroadenoma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e176\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e64.47\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eChronic inflammation\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e2.20\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eIntraductal papilloma\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e7.33\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLobular tumor\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e1.10\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTotal\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e273\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e100.00\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003e2.2 Dataset 2: Validation Set\u003c/h2\u003e\n\u003cp\u003eSimultaneously, 53 lesions from 53 patients were included as Dataset 2, using the same MRI scanner as Dataset 1, but unseen during training. Dataset 2 comprised three subsections: BI-RADS 3, 4, and 5 (see Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Surgical pathology confirmed final diagnoses in cases where percutaneous biopsy indicated high risk. Absence of surgery with imaging stability was deemed indicative of no associated cancer. Follow-up adhered to referenced criteria[\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e]. Correct classification of a lesion required accurate classification in six out of ten images. Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e lists the specific details of Dataset 2.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003ePatients in dataset 2\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eCategory\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eN\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003eConfirmed\u003c/p\u003e\n\u003c/th\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003efollow up\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eB\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eM\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003epathologically\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBI-RADS 3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e16\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBI-RADS 4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBI-RADS 5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003e2.3 MRI Techniques\u003c/h2\u003e\n\u003cp\u003eWe employed two 3T MRI scanners with dedicated breast coils in a prone position. Gd-DTPA (0.1 mmol/kg, 2.50 mL/s) was injected through the elbow vein. The process involved six dynamic enhancement phases (one pre-contrast, five post-contrast). MRIs were conducted preoperatively and before initiating therapy. Detailed scanning parameters are outlined in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eScan parameters for the two magnetic resonance scanners\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eParameter\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePhilips Achieva\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGE Healthcare\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eField strength\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.0T\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.0T\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNo. of coil channels\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAcquisition plane\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAxial\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAxial\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePulse sequence\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3D gradient echo (Thrive)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eEnhanced fast gradient echo 3D\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRepetition time (ms)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5.5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9.6\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eEcho time (ms)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.1\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFlip angle\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10\u0026deg;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10\u0026deg;\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNo. of postcontrast sequence\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFat suppression\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eYes\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eYes\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eScan time\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e570s\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e500s\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003ctfoot\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"3\"\u003e3D, three dimensional; ms, millisecond; s: second.\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tfoot\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003e2.4 Readers\u003c/h2\u003e\n\u003cp\u003eFive experienced radiologists from our department, each with over five years of breast MRI interpretation experience and specialized training in breast imaging, were enlisted. MRI image analyses were conducted using the GOLDPACS viewer (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"http://www.jinpacs.com\" target=\"_blank\"\u003ewww.jinpacs.com\u003c/a\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003e2.5 Proposed model\u003c/h2\u003e\n\u003cp\u003eThe study utilized a computer equipped with an Intel (R) Core (TM) i7-10700F, NVIDIA RTX 2060 GPU, running on Windows 10 Enterprise 64-bit with 6GB RAM. All extraneous programs were closed during model operation. Each network underwent identical data testing and training for consistent comparison. Malignant images were identified based on a threshold of \u0026ge;\u0026thinsp;0.5, while images below this threshold were considered benign.\u003c/p\u003e\n\u003cp\u003eWe selected five commonly used pretrained models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) and employed five-fold cross-validation to assess model performance, selecting the best-performing model. This cross-validation process was then applied to Dataset 2. Additionally, we enhanced model performance using various fine-tuning strategies. The architecture of the proposed DTL with the five models for breast lesion classification is depicted in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eInitially, the images underwent random shuffling. Data augmentation techniques (rotation, shear range, zoom range, and horizontal flip) were applied prior to training. The binary cross-entropy loss function was used, and the training process was optimized using the Adam optimizer with a learning rate of 0.001. Our model required 200 epochs for training on DCE-BMRI images, with a batch size of 64 images. Activation functions included ReLU and sigmoid, as detailed in Equations \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e\u003c/p\u003e\n\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ1\" class=\"mathdisplay\"\u003e$$\\text{R}\\text{e}\\text{l}\\text{u}\\left(\\text{x}\\right)=\\text{f}\\left(\\text{x}\\right)=\\left\\{\\begin{array}{c}max(0,x), x\\ge 0\\\\ 0, x\u0026lt;0\\end{array}\\right.$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ2\" class=\"mathdisplay\"\u003e$$\\text{S}\\text{i}\\text{g}\\text{m}\\text{o}\\text{i}\\text{d}\\left(\\text{x}\\right)=\\text{f}\\left(\\text{x}\\right)=\\frac{1}{1+{\\text{e}}^{-\\text{x}}}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n\u003ch2\u003e2.6 Evaluation metrics\u003c/h2\u003e\n\u003cp\u003eWe assessed the effectiveness of Deep Transfer Learning (DTL) models using five performance metrics: accuracy (Ac), precision (Pr), recall rate (Rc), F1 score (F1), and the area under the receiver operating characteristic curve (AUROC)[\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e]. For this analysis, cases were classified as either malignant or benign, representing positive and negative cases, respectively. True positives (TP) and true negatives (TN) denote the proportion of correctly diagnosed malignant and benign cases. False positives (FP) and false negatives (FN) indicate lesions misdiagnosed as benign and malignant, respectively. The formulas for these metrics are as follows:\u003c/p\u003e\n\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ3\" class=\"mathdisplay\"\u003e$$\\text{A}\\text{c}=\\frac{\\text{T}\\text{P}+\\text{T}\\text{N}}{\\text{T}\\text{P}+\\text{T}\\text{N}+\\text{F}\\text{P}+\\text{F}\\text{N}}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ4\" class=\"mathdisplay\"\u003e$$\\text{P}\\text{r}=\\frac{\\text{T}\\text{P}}{\\text{T}\\text{P}+\\text{F}\\text{P}}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ5\" class=\"mathdisplay\"\u003e$$\\text{R}\\text{c}=\\frac{\\text{T}\\text{P}}{\\text{T}\\text{P}+\\text{F}\\text{N}}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ6\" class=\"mathdisplay\"\u003e$$\\text{F}1=\\frac{2\\times \\text{P}\\text{r}\\times \\text{R}\\text{c}}{\\text{P}\\text{r}+\\text{R}\\text{c}}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eNotably, the accuracy metric (Ac) does not account for data distribution. The F1 score is a balanced measure that considers both precision and recall, making it particularly useful in datasets with imbalanced classes.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n\u003ch2\u003e2.7 Statistical analysis\u003c/h2\u003e\n\u003cp\u003eStatistical analyses were conducted using SPSS 23.0 software (IBM). For data adhering to a normal distribution, counting data were presented as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\stackrel{-}{\\text{x}}\\)\u003c/span\u003e\u003c/span\u003e \u0026plusmn; s). One-way analysis of variance (ANOVA) was employed for variance analysis between groups. The Mann-Whitney U test was applied for data not meeting the normal distribution criteria. The chi-square test was utilized for comparing frequency counts between malignant and benign groups in the datasets (training and testing sets). A P-value of \u0026lt;\u0026thinsp;0.05 (two-tailed) was considered statistically significant.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"3 Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n\u003ch2\u003e3.1 Age and Lesion Diameter\u003c/h2\u003e\n\u003cp\u003eAge and lesion diameter did not conform to normal distribution. The age difference between the malignant group (46.40\u0026thinsp;\u0026plusmn;\u0026thinsp;10.90 years) and the benign group (44.84\u0026thinsp;\u0026plusmn;\u0026thinsp;10.20 years) was not statistically significant (P\u0026thinsp;=\u0026thinsp;0.136). However, lesion diameters were significantly smaller in the malignant group (25.06\u0026thinsp;\u0026plusmn;\u0026thinsp;11.54 mm) compared to the benign group (33.44\u0026thinsp;\u0026plusmn;\u0026thinsp;16.69 mm) (P\u0026thinsp;\u0026lt;\u0026thinsp;0.001). No significant variance was observed in lesion distribution between the training and testing sets across both groups (P\u0026thinsp;=\u0026thinsp;0.988).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003e3.2 Cross validation\u003c/h2\u003e\n\u003cp\u003eWe evaluated five models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) through five-fold cross-validation in Dataset 1 (see Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e for results). The DenseNet201 and MobileNetV2 models achieved perfect accuracy (1.00) in the training set, but their testing set accuracies were lower at 0.91 and 0.88, respectively, both below VGG19\u0026rsquo;s 0.96. Despite similar architectures, VGG19 outperformed VGG16 (0.91). However, both VGG16 and VGG19 exhibited premature loss increases with epoch advancement, indicating non-convergence on Dataset 1 and potential overfitting. Similar trends were observed for MobileNetV2 and DenseNet201. ResNet50 showed the lowest accuracy among the models (0.92 training, 0.67 testing). Figures\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e illustrate the learning curves and heat maps.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab4\" style=\"width: 570px;\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eThe results of the five-fold cross-validation\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth style=\"width: 32.6875px;\" rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eFolds\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 224.312px;\" colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003eAccuracies of the training set\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 225px;\" colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003eAccuracies of the testing set\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003cth style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003emodel1\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel2\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel3\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel4\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel5\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel1\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel2\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel3\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel4\u003c/p\u003e\n\u003c/th\u003e\n\u003cth style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003emodel5\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 32.6875px;\" align=\"left\"\u003e\n\u003cp\u003eFold1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 32.6875px;\" align=\"left\"\u003e\n\u003cp\u003eFold2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.99\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 32.6875px;\" align=\"left\"\u003e\n\u003cp\u003eFold3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 32.6875px;\" align=\"left\"\u003e\n\u003cp\u003eFold4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.99\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 32.6875px;\" align=\"left\"\u003e\n\u003cp\u003eFold5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 44.3125px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.96\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 45px;\" align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNote. Model1, VGG16; Model2, VGG19; Model3, DenseNet201; Model4, ResNet50; Model5, MobileNetV2.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n\u003ch2\u003e3.3 Fine-tuning strategy\u003c/h2\u003e\n\u003cp\u003eGiven these findings, we focused on enhancing the VGG19 model through five distinct fine-tuning strategies (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). The fine-tuning involved activating neural network parameters for training, while keeping certain layers frozen. We noted that the accuracy achieved was 1.0, for all five fine-tuning models(S1-5) on the training set, but S4 obtained the highest test accuracy of 0.97 on the testing set. In addition, the loss value was the lowest in the testing set for S4. These results reveal that the S4 model has a better generalization ability than the other fine-tuned models.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n\u003ch2\u003e3.4 ROC Analysis on Validation Set\u003c/h2\u003e\n\u003cp\u003eAnalysis of the Receiver Operating Characteristic (ROC) curve for the five models on the validation set revealed VGG19 as the highest performer (AUC 0.92), yet the validation set AUC was only 0.76. Among the fine-tuned models, S4 attained the highest AUC (0.89) on the validation set, marking a 13% improvement over the original VGG19 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e). Further analysis of S4 across BI-RADS categories 3, 4, and 5 showed notably higher AUCs for BI-RADS 3 (0.90) and 4 (0.86) compared to 5 (0.65) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\n\u003ch2\u003e3.5 Classification Reports on Validation Set\u003c/h2\u003e\n\u003cp\u003eClassification reports for the five models and S1-5 strategies are provided in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. For the validation set, VGG19 achieved higher performance metrics (Pr 0.75, Rc 0.76, F1 0.73, AUC 0.76) compared to the other models. Strategy S4 outperformed all others on the validation set with Pr 0.89, Rc 0.88, F1 0.87, and AUC 0.89.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab5\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eClassification report of deep transfer learning models in validation set\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth colspan=\"2\" rowspan=\"3\" align=\"left\"\u003e\n\u003cp\u003eDTL\u003c/p\u003e\n\u003cp\u003emodels\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003ePr\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003eRc\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\text{f}1\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"1\" align=\"left\"\u003eAUC\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eavg\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eavg\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003egroup1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eavg\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emodel1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.53\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.70\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emodel2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.59\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.63\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emodel3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.52\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.59\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.69\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emodel4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.39\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.53\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.84\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.68\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.53\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.65\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emodel5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.55\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.61\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.89\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.68\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eS1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.63\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eS2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.95\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.64\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.94\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eS3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.63\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eS4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.98\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.89\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.78\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.98\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.89\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eS5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.94\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.59\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.64\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNote. Group1, benign group; Group2, malignant group; Avg, average; Model1, VGG16; Model2, VGG19; Model3, DenseNet201; Model4, ResNet50; Model5, MobileNetV2.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e"},{"header":"4 Discussion","content":"\u003cp\u003eIn this study, we evaluated five pre-trained convolutional neural network models using a 5-fold cross-validation method on our DCE-BMRI dataset. Our objective was to identify the best-performing model, which we defined as the one excelling across all predefined evaluation criteria. Following this, we fine-tuned the chosen model to enhance its performance further and tested its generalization capability on a validation set.\u003c/p\u003e \u003cp\u003eOur findings indicated that the VGG19 model demonstrated superior performance, achieving accuracies of 1.00 and 0.96 on the training and testing sets, respectively. Moreover, VGG19 achieved the highest Area Under the Curve (AUC) of 0.92 on the first validation set, but this dropped to 0.76 on a subsequent validation set, suggesting limitations in its generalization ability. Previous research supports the notion that fine-tuning can enhance the accuracy and precision of such models[\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eConsequently, we developed five distinct fine-tuning strategies for VGG19. Strategy S4 emerged as the most successful, yielding the highest test accuracy (0.97) and the lowest test loss on the validation set, indicating a superior generalization capability compared to the other strategies. When comparing the AUC scores of strategies S1-5 on the validation set, S4 again scored highest with an AUC of 0.89. These results are promising for advancing the accuracy of medical image classification diagnostics.\u003c/p\u003e \u003cp\u003eWe also delved into whether the S4 model exhibited different AUC scores across BI-RADS categories 3, 4, and 5. Interestingly, S4 performed best in BI-RADS 3 (AUC 0.90), followed by BI-RADS 4 (AUC 0.86), and showed the least performance in BI-RADS 5 (AUC 0.65). The BI-RADS 3 category, typically applied when the likelihood of cancer is less than 2%, aims to minimize unnecessary biopsies for pathologically benign findings. However, patient compliance with follow-up MRI recommendations every six months is notably low in this category[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe BI-RADS system is set to evolve with new breast imaging modalities. Key areas for improvement include expanding the lexicon for common findings and clarifying the application of Category 3[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. BI-RADS 3 represents a significant portion (13.9%) of diagnostic exams [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], often leading to follow-up procedures for patients classified as \u0026lsquo;probably benign\u0026rsquo;, yet compliance with these follow-up recommendations remains a challenge [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. This lack of compliance raises concerns about the clinical and economic implications, particularly regarding the resolution time and outcomes for patients in this category.\u003c/p\u003e \u003cp\u003eIn a particular study, only 1.4% of BI-RADS 3 lesions were found to be malignant, including two cases of delayed diagnosis at 13.2 and 33.2 months, respectively. The incidence of delayed diagnosis due to additional MRI-detected lesions during follow-up was notably low (0.7%), consisting exclusively of T1N0 contralateral cancers. This finding suggests that annual follow-up may suffice for BI-RADS 3 lesions identified by MRI before surgery[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Consequently, accurately distinguishing between benign and malignant lesions in BI-RADS 3 is crucial. Our study potentially offers significant benefits to a substantial number of patients diagnosed with BI-RADS category lesions using DCE-BMRI imaging.\u003c/p\u003e \u003cp\u003eBI-RADS category 4 lesions are associated with a high likelihood of malignancy, with estimates ranging from 2\u0026ndash;95%. The BI-RADS 4 classification, to a degree, is subjective; the outcomes of biopsies in this category vary significantly, and the rate of cancer detection relative to the number of biopsies performed is relatively low (17.8%)[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Moreover, unnecessary biopsies can result in a range of adverse effects, including pain, fear, emotional distress, and financial costs.\u003c/p\u003e \u003cp\u003eBreast MRI is highly sensitive, yet it often presents a challenge in differentiating between atypical malignant and benign lesions, leading to potential overclassification in the BI-RADS 4 category and subsequent invasive biopsies. The wide range of positive predictive values for MRI-guided biopsies (2.5\u0026ndash;84.0%) [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], indicates that many patients undergo unnecessary procedures. Indeed, numerous women subjected to biopsies for benign findings endure unnecessary discomfort, expenses, potential complications, cosmetic alterations, and anxiety [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Identifying predictors of benign BI-RADS 4 masses, therefore, could be highly beneficial.\u003c/p\u003e \u003cp\u003eEfforts to enhance the assessment of BI-RADS 4 lesions could improve the identification of benign lesions, thereby reducing the frequency of unnecessary biopsies. Some researchers have developed predictive models based on imaging features or multiparameter MRI data to better evaluate BI-RADS 4 lesions, though these models typically rely on traditional imaging features subjectively defined by radiologists[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe BI-RADS 5 category is applied when imaging findings suggest a malignancy probability of 95% or higher. According to MC et al.[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], the positive predictive value of BI-RADS 5 assessments is only 71.4%, indicating that not all lesions classified as BI-RADS 5 are malignant[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], and surgery is often recommended for this category. It is recognized that a single imaging finding rarely confers such a high risk of malignancy; rather, a combination of features is necessary to elevate a lesion to Category 5[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. However, it's important to acknowledge that even when tissue samples from molecular biopsies are used, they may not fully represent the entire lesion, as biopsies often target only a small, specific area of a heterogeneous lesion, introducing a bias in lesion selection.\u003c/p\u003e \u003cp\u003eRc also known as sensitivity, measures a classifier's completeness. A lower Rc value indicates the classifier's limited capability in handling large FP values. Recent publications have led to the introduction of new and updated performance benchmarks in the latest edition, replacing outdated metrics. The recall rate benchmark has been revised accordingly. Initially, about half of all radiologists were unable to meet the 10% benchmark for recall rate, prompting a revision to a more achievable target of 12%, a standard met by over 75% of radiologists[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Our study showed that the S4 model exhibited the highest recall rate (0.89) among all DTL models, which is notable given the relatively limited class diversity in our dataset.\u003c/p\u003e \u003cp\u003eThis study, however, is not without limitations. Firstly, the training set contained a relatively small number of images, particularly lacking in rare lesion types. Consequently, our dataset may not fully represent the broader spectrum of breast disease patients, potentially impacting the accuracy of the DTL model. Therefore, it\u0026rsquo;s essential to conduct further analysis with more extensive datasets to comprehensively evaluate the robustness of the DTL model. Secondly, our study relied solely on static Dynamic Contrast-Enhanced Breast Magnetic Resonance Imaging (DCE-BMRI) images, excluding other routine diagnostic procedures such as clinical evaluations, breast ultrasounds, and mammography. Thirdly, we limited our investigation to just five pre-trained models; future research should explore a wider range of models to determine their robustness on larger datasets. Lastly, while this paper does not delve into the various methods of fine-tuning Convolutional Neural Network (CNN) models, these topics will be the focus of our subsequent studies.\u003c/p\u003e"},{"header":"5 Conclusions","content":"\u003cp\u003eThe S4 model demonstrated superior accuracy in BI-RADS categories 3 and 4 compared to category 5. This finding is significant as it could lead to a reduction in the number of follow-up sessions for BI-RADS 3 and decrease the number of unnecessary biopsies for benign lesions in BI-RADS 4. However, these results necessitate further validation. Moving forward, our focus will be on exploring more robust models and the importance of augmenting our dataset with additional data.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eMRI= Magnetic resonance imaging\u003c/p\u003e\n\u003cp\u003eDL= Deep learning\u003c/p\u003e\n\u003cp\u003eDTL=Deep transfer learning\u003c/p\u003e\n\u003cp\u003eROC = receiver operating characteristic\u003c/p\u003e\n\u003cp\u003eAUC = area under the ROC curve\u003c/p\u003e\n\u003cp\u003eBI-RADS = Breast Imaging Reporting and Data System,\u003c/p\u003e\n\u003cp\u003eDCE-BMRI = dynamic contrast enhanced breast MRI\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was approved by the Second Hospital of Changzhou Affiliated to Nanjing Medical University of Chinese Medicine Ethics Review Committee (Ethics Number: [2023]KY313-01).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was a retrospective analysis and informed consent was waived.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated during and/or analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported by the Program of Bureau of Science and Technology Foundation of Changzhou (No. CJ20220260).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLi Li and Mingzhu Meng carried out the literature search, and designed and wrote the manuscript. Mingzhu Meng and Changjie Pan conceived of the project, and participated in its design and coordination, and helped to draft the manuscript. Ming Zhang, Dong Shen and Guangyuan He are responsible for figures processing. Both authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors wish to thank Shiquan Ge for his technical assistance in operating the Python programming code.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSung H, Ferlay J, Siegel RL,et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209-49.\u003c/li\u003e\n\u003cli\u003eDeSantis CE, Ma J, Gaudet MM,et al. Breast cancer statistics, 2019. CA Cancer J Clin. 2019;69(6):438-51.\u003c/li\u003e\n\u003cli\u003eCao W, Chen HD, Yu YW,et al. Changing profiles of cancer burden worldwide and in China: a secondary analysis of the global cancer statistics 2020. Chin Med J (Engl). 2021;134(7):783-91.\u003c/li\u003e\n\u003cli\u003eRM F, YN Z, SM C,et al. Current cancer situation in China: good or bad news from the 2018 Global Cancer Statistics? Cancer Commun (Lond). 2019;39(1):22.\u003c/li\u003e\n\u003cli\u003eGao Y, Heller SL. Abbreviated and Ultrafast Breast MRI in Clinical Practice. RadioGraphics. 2020;40(6):1507-27.\u003c/li\u003e\n\u003cli\u003eBerdzuli N. Breast cancer: from awareness to access. BMJ. 2023;380(290.\u003c/li\u003e\n\u003cli\u003eMercado CL. BI-RADS Update. Radiol Clin North Am. 2014;52(3):481-7.\u003c/li\u003e\n\u003cli\u003eSedgwick EL, Ebuoma L, Hamame A,et al. BI-RADS update for breast cancer caregivers. Breast Cancer Research and Treatment. 2015;150(2):243-54.\u003c/li\u003e\n\u003cli\u003ePesce K, Orruma MB, Hadad C,et al. BI-RADS Terminology for Mammography Reports: What Residents Need to Know. Radiographics. 2019;39(2):319-20.\u003c/li\u003e\n\u003cli\u003eHuang Z, Zhou Q, Zhu X,et al. Batch Similarity Based Triplet Loss Assembled into Light-Weighted Convolutional Neural Networks for Medical Image Classification. Sensors (Basel). 2021;21(3):764-85.\u003c/li\u003e\n\u003cli\u003eTajbakhsh N, Jeyaseelan L, Li Q,et al. Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation. Medical image analysis. 2020;63:101693.\u003c/li\u003e\n\u003cli\u003eNaqa IE. The role of machine and deep learning in modern medical physics. Med Phys. 2020;47(5):e125-e6.\u003c/li\u003e\n\u003cli\u003eZhang J, Xie Y, Wu Q,et al. Medical image classification using synergic deep learning. Medical image analysis. 2019;54:10-9.\u003c/li\u003e\n\u003cli\u003eDong F, She R, Cui C,et al. One step further into the blackbox: a pilot study of how to build more confidence around an AI-based decision system of breast nodule assessment in 2D ultrasound. European radiology. 2021;31(7):4991-5000.\u003c/li\u003e\n\u003cli\u003eLee SE, Lee JH, Han K,et al. BI-RADS category 3, 4, and 5 lesions identified at preoperative breast MRI in patients with breast cancer: implications for management. Eur Radiol Exp. 2020;30(5):2773-81.\u003c/li\u003e\n\u003cli\u003eWang Z, Li X, Yao M,et al. A new detection model of microaneurysms based on improved FC-DenseNet. Sci Rep. 2022;12(1):950.\u003c/li\u003e\n\u003cli\u003eTan T, Li Z, Liu H,et al. Optimize Transfer Learning for Lung Diseases in Bronchoscopy Using a New Concept: Sequential Fine-Tuning. IEEE J Transl Eng Health Med. 2018;6:1800808.\u003c/li\u003e\n\u003cli\u003eAhamed KU, Islam M, Uddin A,et al. A deep learning approach using effective preprocessing techniques to detect COVID-19 from chest CT-scan and X-ray images. Comput Biol Med. 2021;139:105014.\u003c/li\u003e\n\u003cli\u003eMontaha S, Azam S, Rafid A,et al. BreastNet18: A High Accuracy Fine-Tuned VGG16 Model Evaluated Using Ablation Study for Diagnosing Breast Cancer from Enhanced Mammography Images. Biology (Basel). 2021;10(12):1347.\u003c/li\u003e\n\u003cli\u003eAmbinder EB, Myers K, Panigrahi B,et al. Breast MRI BI-RADS 3: Impact of Patient-Level Factors on Compliance With Short-Term Follow-Up. J Am Coll Radiol 2020;17(3):377-83.\u003c/li\u003e\n\u003cli\u003eEghtedari M, Chong A, Rakow-Penner R,et al. Current Status and Future of BI-RADS in Multimodality Breast Imaging, From the AJR Special Series on Radiology Reporting and Data Systems. American Journal of Roentgenology. 2020:[published online].\u003c/li\u003e\n\u003cli\u003eElezaby M, Li G, Bhargavan-Chatfield M,et al. ACR BI-RADS assessment category 4 subdivisions in Diagnostic Mammography: Utilization and Outcomes in the National Mammography Database. Radiology. 2018;287(2):416-22.\u003c/li\u003e\n\u003cli\u003eRM S, ES B, M E,et al. Utility of BI-RADS Assessment Category 4 Subdivisions for Screening Breast MRI. AJR American journal of roentgenology. 2017;208(6):1392-9.\u003c/li\u003e\n\u003cli\u003eJR MdA, AB G, TP B,et al. Subcategorization of Suspicious Breast Lesions (BI-RADS Category 4) According to MRI Criteria: Role of Dynamic Contrast-Enhanced and Diffusion-Weighted Imaging. AJR American journal of roentgenology. 2015;205(1):222-31.\u003c/li\u003e\n\u003cli\u003eLaws A, Crocker A, Dort J,et al. Improving Wait Times and Patient Experience Through Implementation of a Provincial Expedited Diagnostic Pathway for BI-RADS 5 Breast Lesions. Ann Surg Oncol 2019;36(10):3361-7.\u003c/li\u003e\n\u003cli\u003eHao W, Gong J, Wang S,et al. Application of MRI Radiomics-Based Machine Learning Model to Improve Contralateral BI-RADS 4 Lesion Assessment. Frontiers in Oncology. 2020;10:531476-84.\u003c/li\u003e\n\u003cli\u003eMC M, C G, L H,et al. Positive Predictive Value of BI-RADS MR imaging. Radiology. 2012;264(1):51-8.\u003c/li\u003e\n\u003cli\u003eDao KA, Rives AF, Quintana LM,et al. BI-RADS 5: More than Cancer. Radiographics. 2020;40(5):1203-4.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-imaging","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmim","sideBox":"Learn more about [BMC Medical Imaging](http://bmcmedimaging.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmim/default.aspx","title":"BMC Medical Imaging","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"BI-RADS, Convolutional Neural Networks, Deep transfer learning, Breast lesions, Magnetic resonance imaging","lastPublishedDoi":"10.21203/rs.3.rs-3937557/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3937557/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjectives.\u003c/h2\u003e \u003cp\u003eThis study aimed to explore which convolutional neural network (CNN) model is best for predicting the likelihood of malignancy on dynamic contrast-enhanced breast magnetic resonance imaging (DCE-BMRI).\u003c/p\u003e\u003ch2\u003eMaterials and Methods.\u003c/h2\u003e \u003cp\u003eA total of 273 benign (benign group) and 274 malignant lesions (malignant group) were obtained, and randomly divided into a training set (benign group: 246 lesions, malignant group: 245 lesions) and a testing set (benign group: 28 lesions, malignant group: 28 lesions) in a 9:1 ratio. An additional 53 lesions from 53 patients were designated as the validation set. Five models (VGG16, VGG19, DenseNet201, ResNet50, and MobileNetV2) were evaluated. The metrics for model performance evaluation included accuracy (Ac) in the training and testing sets, and precision (Pr), recall rate (Rc), F1 score (F1), and area under the receiver operating characteristic curve (AUC) in the validation set.\u003c/p\u003e\u003ch2\u003eResults.\u003c/h2\u003e \u003cp\u003eAccuracies of 1.0 were achieved on the training set by all five fine-tuned models (S1-5), with model S4 demonstrating the highest test accuracy at 0.97. Additionally, S4 showed the lowest loss value in the testing set. The S4 model also attained the highest AUC (Area Under the Curve) of 0.89 in the validation set, marking a 13% improvement over the VGG19 model. Notably, the AUC of S4 for BI-RADS 3 was 0.90 and for BI-RADS 4 was 0.86, both significantly higher than the 0.65 AUC for BI-RADS 5.\u003c/p\u003e\u003ch2\u003eConclusion.\u003c/h2\u003e \u003cp\u003eThe S4 model we propose emerged as the superior model for predicting the likelihood of malignancy in DCE-BMRI and holds potential for clinical application in patients with breast diseases. However, further validation is necessary, underscoring the need for additional data.\u003c/p\u003e","manuscriptTitle":"Predicting Malignancy in Breast Lesions: Enhancing Accuracy with Fine-Tuned Convolutional Neural Network Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-20 18:00:05","doi":"10.21203/rs.3.rs-3937557/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-08-19T10:14:30+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-26T07:46:27+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-21T10:49:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"102518760829480346860913456933157230161","date":"2024-07-19T13:54:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"187041833816270480231627125467668471846","date":"2024-07-19T09:31:46+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-04-19T15:43:10+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-02-16T04:41:59+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-02-16T04:39:47+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-02-16T04:39:47+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Imaging","date":"2024-02-07T16:58:11+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-imaging","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmim","sideBox":"Learn more about [BMC Medical Imaging](http://bmcmedimaging.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmim/default.aspx","title":"BMC Medical Imaging","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2d2cfdca-c19e-4a18-a2c8-206eaf6d17eb","owner":[],"postedDate":"February 20th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-11-18T19:16:51+00:00","versionOfRecord":{"articleIdentity":"rs-3937557","link":"https://doi.org/10.1186/s12880-024-01484-1","journal":{"identity":"bmc-medical-imaging","isVorOnly":false,"title":"BMC Medical Imaging"},"publishedOn":"2024-11-11 15:58:13","publishedOnDateReadable":"November 11th, 2024"},"versionCreatedAt":"2024-02-20 18:00:05","video":"","vorDoi":"10.1186/s12880-024-01484-1","vorDoiUrl":"https://doi.org/10.1186/s12880-024-01484-1","workflowStages":[]},"version":"v1","identity":"rs-3937557","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3937557","identity":"rs-3937557","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.