Pediatric Appendicitis Detection from Ultrasound Images | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Pediatric Appendicitis Detection from Ultrasound Images Fatemeh Hosseinabadi, Seyedhassan Sharifi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7866377/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Pediatric appendicitis remains one of the most common causes of acute abdominal pain in children, and its diagnosis continues to challenge clinicians due to overlapping symptoms and variable imaging quality. This study aims to develop and evaluate a deep learning model based on a pretrained ResNet architecture for automated detection of appendicitis from B-mode ultrasound images. We used the Regensburg Pediatric Appendicitis Dataset, which includes ultrasound scans, laboratory data, and clinical scores from pediatric patients admitted with abdominal pain to Children’s Hospital St. Hedwig in Regensburg, Germany (2016–2021). Each subject had 1–15 ultrasound views covering the right lower quadrant, appendix, lymph nodes, and related structures. For the image-based classification task, ResNet was fine-tuned to distinguish appendicitis from non-appendicitis cases. Images were preprocessed by normalization, resizing, and augmentation to enhance generalization. The proposed ResNet model achieved an overall accuracy of 93.44%, precision of 91.53%, and recall of 89.8%, demonstrating strong performance in identifying appendicitis across heterogeneous ultrasound views. The model effectively learned discriminative spatial features, overcoming challenges posed by low contrast, speckle noise, and anatomical variability in pediatric imaging. Pediatric Appendicitis Ultrasound ResNet Figures Figure 1 1. Introduction Acute appendicitis is the most common surgical emergency in children and adolescents, accounting for approximately 1–2 cases per 1,000 individuals annually. It typically results from luminal obstruction of the appendix, leading to inflammation, bacterial overgrowth, and eventual perforation if left untreated. The condition can progress rapidly from uncomplicated to complicated appendicitis, potentially resulting in peritonitis, abscess formation, and life-threatening sepsis. In pediatric patients, early and accurate diagnosis is essential because younger children often present with atypical symptoms, making clinical differentiation from other causes of abdominal pain more challenging. Studies have reported that diagnostic errors in appendicitis remain a significant concern, with misdiagnosis rates ranging from 15% to 30% in some age groups. False-negative diagnoses increase the risk of complications, while false-positive diagnoses may lead to unnecessary appendectomies, prolonged hospitalization, and higher medical costs. Consequently, improving diagnostic precision is a priority in pediatric emergency medicine [ 1 , 2 ]. Current diagnostic workflows for suspected appendicitis combine clinical evaluation, laboratory testing, and imaging. Clinical scoring systems, such as the Alvarado Score and Pediatric Appendicitis Score (PAS), integrate symptoms (e.g., pain migration, nausea, tenderness) and laboratory markers (e.g., leukocytosis, elevated C-reactive protein) [ 3 ]. Although these scores are helpful for risk stratification, they are not definitive and often require confirmation through imaging studies. Among imaging modalities, ultrasound (US) is the preferred first-line technique for children due to its noninvasive nature, lack of radiation, and availability in emergency settings. However, ultrasound diagnosis of appendicitis is highly operator-dependent and subject to variability in patient anatomy, bowel gas interference, and body composition. Visualization of the appendix is successful in only 60–80% of cases, and diagnostic accuracy can drop significantly in obese or uncooperative children. Even experienced radiologists may encounter difficulty differentiating early appendicitis from mesenteric lymphadenitis or gastrointestinal infections, particularly when image quality is suboptimal. These limitations create an urgent need for automated image interpretation systems capable of assisting clinicians with consistent and objective diagnostic insights [ 4 , 5 ]. In recent years, artificial intelligence (AI) and machine learning (ML) have revolutionized diagnostic imaging by allowing computational models to identify complex visual and statistical patterns beyond human perception [ 6 – 9 ]. In radiology, AI has shown significant potential in tasks such as tumor detection in MRI and CT, lung pathology screening in chest X-rays, and cardiac function assessment in echocardiography [ 10 , 11 ]. For ultrasound imaging specifically, AI algorithms have been successfully applied to fetal growth monitoring, thyroid nodule classification, and liver fibrosis staging [ 12 – 15 ]. The main advantage of AI-driven systems lies in their ability to learn directly from raw image data, thereby minimizing dependence on subjective interpretations. Deep learning, particularly Convolutional Neural Networks (CNNs), has emerged as the cornerstone of medical image analysis due to its ability to automatically extract hierarchical features from low-level edges and textures to high-level structural patterns that correspond to anatomical and pathological cues. This makes CNNs particularly suitable for ultrasound data, where signal-to-noise ratios are low, and manual feature engineering is often inadequate. Among various CNN architectures, Residual Networks (ResNets) have demonstrated outstanding performance in both general computer vision and medical image analysis. ResNets introduce shortcut (skip) connections that bypass one or more layers, allowing the network to learn residual mappings instead of direct transformations. This study aims to harness the power of deep residual learning to improve the diagnostic accuracy of pediatric appendicitis detection using ultrasound data. We employed the Regensburg Pediatric Appendicitis Dataset, a comprehensive dataset collected from pediatric patients admitted with abdominal pain between 2016 and 2021 at the Children’s Hospital St. Hedwig in Regensburg, Germany. The dataset includes B-mode ultrasound images, laboratory findings, clinical scores, and expert annotations. Our primary goal was to train a ResNet-based CNN to classify ultrasound images as appendicitis or non-appendicitis and to assess its diagnostic performance against standard metrics. 2. Method 2.1 Dataset Description This study utilized the Regensburg Pediatric Appendicitis Dataset, a curated clinical dataset collected retrospectively from pediatric patients admitted with abdominal pain to Children’s Hospital St. Hedwig, Regensburg, Germany, between 2016 and 2021. The dataset includes a rich combination of imaging, clinical, and laboratory information to support multimodal diagnostic modeling. Each patient record may contain one to fifteen B-mode ultrasound (US) images, captured from multiple abdominal regions of interest such as the right lower quadrant (RLQ), appendix, intestinal loops, lymph nodes, free fluid areas, and reproductive organs. The ultrasound images are stored in BMP format under the US_Pictures/ directory, with filenames corresponding to subject identifiers and view indices (e.g., 23.7.bmp for patient 23, view 7) [ 16 , 17 ]. In addition to imaging data, the accompanying file app_data.xlsx contains tabular variables summarizing laboratory test results, physical examination findings, and expert ultrasonographic assessments. Clinical scoring systems such as the Alvarado Score and Pediatric Appendicitis Score (PAS) were included to contextualize imaging results with established diagnostic criteria. Each subject is annotated for three outcome variables: Diagnosis – appendicitis vs. no appendicitis, Management – surgical vs. conservative treatment, and Severity – complicated vs. uncomplicated (or no appendicitis). The study was approved by the Ethics Committee of the University of Regensburg (no. 18-1063- 101, 18-1063_1-101 and 18-1063_2-101) and was performed following applicable guidelines and regulations. The ethics committee confirmed that there was no need for written informed consent for the retrospective analysis and publication of anonymized routine data according to Art. 27 para. 4 of the Bavarian Hospital Law. For patients followed up after discharge, written informed consent was obtained from parents or legal representatives. In this work, only the Diagnosis label (appendicitis vs. no appendicitis) was used for binary image classification. 2.2 Data Preprocessing The ultrasound images exhibited high inter-patient variability in acquisition parameters such as brightness, contrast, and spatial scale. To standardize inputs, each image was converted to grayscale, resized to 224 × 224 pixels, and normalized to zero mean and unit variance. Data augmentation was applied to improve model generalization and simulate clinical variability, including random rotations (± 10°), horizontal flips, contrast adjustments, and Gaussian noise injection. Because multiple images were available per subject, all images were treated as independent samples while ensuring that images from the same patient were confined to either the training or testing split to avoid data leakage. The final dataset was divided into 80% for training and 20% for testing, maintaining balanced class proportions. 2.3 Model Architecture We implemented a Convolutional Neural Network (CNN) based on ResNet architecture, leveraging transfer learning to benefit from pre-learned visual representations. Specifically, ResNet-50 pretrained on the ImageNet dataset was fine-tuned for the ultrasound classification task. The network structure consisted of: An input layer receiving 224 × 224 × 3 normalized images. The initial convolution and max-pooling layers from the base ResNet-50 model. Four residual blocks, each containing multiple convolutional layers with skip connections that enable residual learning and mitigate vanishing gradients. Global average pooling to condense spatial information. A fully connected dense layer with ReLU activation for feature integration. A final sigmoid output layer producing probabilities for the binary classes (appendicitis vs. no appendicitis). During fine-tuning, the earlier layers were frozen to preserve general low-level features, while the later residual blocks and fully connected layers were retrained to adapt to ultrasound-specific texture patterns. 2.4 Training Procedure Model implementation and training were conducted in Python 3.10 using TensorFlow 2.14 / Keras on an NVIDIA GPU workstation. Training used the following hyperparameters: Optimizer: Adam (learning rate = 1 × 10⁻⁴) Loss function: Binary cross-entropy Batch size: 32 Epochs: 100 (early stopping based on validation loss) Dropout: 0.3 on the dense layer to prevent overfitting Each training epoch included on-the-fly data augmentation. Model checkpoints and validation metrics were recorded at each epoch. Fine-tuning the upper residual blocks improved convergence and boosted overall classification accuracy. 2.5 Evaluation Metrics Model performance was evaluated using standard classification metrics: accuracy (ACC), precision (PRE), recall (REC), F1-score, and area under the receiver-operating characteristic curve (AUC). 3. Results 3.1 Model Performance The proposed ResNet-based deep learning model demonstrated strong performance in detecting pediatric appendicitis from ultrasound images. After fine-tuning the pretrained ResNet-50 architecture, the model achieved an overall accuracy of 93.44%, a precision of 91.53%, and a recall (sensitivity) of 89.8% on the held-out test dataset. The F1-score, representing the balance between precision and recall, was calculated at 90.6%, indicating a stable and reliable detection capability. The Receiver Operating Characteristic (ROC) curve exhibited an Area Under the Curve (AUC) of 0.95, reflecting excellent discriminative power between appendicitis and non-appendicitis cases. The confusion matrix (Figure X) illustrated that most appendicitis cases were correctly identified, with only a small number of false negatives, primarily in borderline or low-quality ultrasound images. False positives were predominantly associated with cases showing inflamed lymph nodes or bowel wall thickening, which can mimic appendicitis sonographically. Nevertheless, the model’s high precision underscores its ability to minimize false alarms and provide radiologists with reliable assistance in triaging ambiguous cases. 3.2 Training and Validation Curves Figure X presents the training and validation accuracy and loss curves over 100 epochs. The training process exhibited steady convergence, with both training and validation accuracy improving consistently without significant overfitting. Early stopping based on validation loss prevented degradation of generalization performance. The final validation loss stabilized at 0.184, suggesting that the model effectively learned meaningful features without memorizing noise or irrelevant textures from the ultrasound data. The inclusion of dropout regularization and data augmentation contributed to stable training behavior and robust generalization. 3.3 Visual Feature Interpretation Feature activation maps generated from the final convolutional layers using Gradient-weighted Class Activation Mapping (Grad-CAM) provided qualitative insights into model interpretability (Figure X). The heatmaps revealed that the network consistently focused on anatomically relevant regions such as the appendiceal area, pericecal fat, and surrounding bowel loops, aligning well with radiologists’ regions of interest during manual assessment. This correspondence between AI attention and clinical focus reinforces the physiological relevance of the learned representations, suggesting that the ResNet model’s predictions are grounded in meaningful image features rather than artifacts or background textures. In several correctly classified appendicitis cases, Grad-CAM visualizations highlighted inflamed tubular structures and peri-appendiceal fat echogenicity, both key sonographic indicators of appendiceal inflammation. In non-appendicitis cases, the model concentrated on other abdominal regions, confirming the absence of the pathological pattern. Such visualization tools enhance model transparency and can aid radiologists in understanding the reasoning behind automated classifications. 3.4 Comparison with Previous Approaches Previous research on appendicitis detection using traditional machine learning methods relied primarily on handcrafted features, such as gray-level co-occurrence matrices (GLCM), edge descriptors, and statistical intensity distributions, combined with classifiers like Support Vector Machines (SVMs) or Random Forests. Reported accuracies in these methods typically ranged from 75% to 85%, limited by the subjectivity of feature engineering and the inherent variability of ultrasound image quality. In contrast, the proposed ResNet-based deep learning approach automatically extracted hierarchical spatial features directly from the ultrasound data, eliminating the need for manual feature design. This end-to-end learning framework not only improved accuracy to 93.44% but also offered superior robustness to noise, variable acquisition settings, and anatomical diversity. Furthermore, the use of transfer learning significantly reduced the amount of required labeled data and training time compared to models trained from scratch. Table X summarizes the quantitative performance of the proposed model. The close alignment between accuracy, precision, and recall indicates that the model maintains balanced classification performance across both classes, avoiding bias toward either appendicitis or normal samples. 3.5 Clinical Relevance From a clinical standpoint, the model’s high sensitivity (recall) is particularly valuable in reducing missed appendicitis cases, which can lead to severe complications if untreated. Likewise, the strong precision reduces the likelihood of false-positive diagnoses, which may otherwise result in unnecessary imaging or surgical intervention. Integrating such AI tools into clinical workflows could assist radiologists, especially in resource-limited or high-volume settings, by providing real-time decision support and standardized interpretation across operators. Table 1 Classification performance metric Metric Value (%) Accuracy 93.44 Precision 91.53 Recall (Sensitivity) 89.80 F1-Score 90.6 AUC 95.0 4. Discussion 4.1 Summary of Findings This study developed and validated a deep residual convolutional neural network (ResNet-50) for the automatic detection of pediatric appendicitis using B-mode ultrasound images from the Regensburg Pediatric Appendicitis Dataset. The model achieved an overall accuracy of 93.44%, with a precision of 91.53% and a recall of 89.8%, demonstrating that a deep learning framework can accurately identify appendicitis in children using noninvasive imaging data. These findings highlight the potential of AI-assisted diagnostic tools to complement radiologist interpretations, particularly in emergency and resource-constrained clinical environments where rapid, objective, and reproducible results are essential. The high accuracy achieved in this study surpasses the performance reported in many traditional machine learning approaches, which often rely on handcrafted features extracted from ultrasound intensity patterns or textural statistics. By contrast, the proposed ResNet model automatically learned spatially and contextually rich representations directly from imaging data, effectively capturing the structural and morphological characteristics of the inflamed appendix and surrounding tissues. 4.2 Comparison with Previous Studies Previous research efforts in automated appendicitis diagnosis have explored various imaging modalities, including CT, MRI, and ultrasound, with machine learning models such as support vector machines (SVMs), k-nearest neighbors (k-NN), and random forests. For example, studies using CT-based deep learning classifiers reported accuracies between 85% and 92%, albeit at the cost of radiation exposure — a significant drawback for pediatric populations. Other ultrasound-based studies employing traditional ML approaches achieved performance typically below 85% due to the limited generalizability of manually engineered features. Our findings align with recent advances in deep learning for pediatric imaging, where transfer learning using pretrained CNN architectures has shown notable improvements in diagnostic performance. The ResNet model used in this study leverages residual learning, which enables the network to train deeper architectures without the risk of gradient degradation. This design allows the model to learn both low-level ultrasound textures and high-level semantic representations critical for discriminating appendicitis from other abdominal conditions. Moreover, Grad-CAM visualization confirmed that the model’s focus regions overlapped with clinically relevant anatomical sites, lending interpretability and biological credibility to the predictions. 4.3 Clinical Implications Accurate diagnosis of pediatric appendicitis remains a persistent challenge, as clinical symptoms are often nonspecific and imaging results may be inconclusive. The proposed AI-driven approach has the potential to augment radiologist performance by providing a rapid, consistent, and objective assessment of ultrasound images. In emergency departments, such models could serve as second readers, flagging suspicious cases for further evaluation and helping to standardize diagnostic decisions across varying levels of clinical expertise. Importantly, this system operates entirely on noninvasive ultrasound imaging, which is safer for pediatric patients than CT-based protocols. The integration of such deep learning systems into clinical decision support platforms could reduce the diagnostic delay and variability that currently affect appendicitis management. For instance, early AI-assisted identification of appendicitis could enable faster surgical consultations, minimize unnecessary hospital admissions, and optimize the use of imaging resources. Ultimately, these tools may contribute to lowering rates of perforation and postoperative complications by facilitating timely and accurate diagnosis. 4.4 Interpretability and Trust in AI Models One major barrier to clinical adoption of AI systems is the lack of interpretability. Deep learning models are often viewed as “black boxes,” which can reduce clinician trust in automated outputs. To address this, the present study incorporated visual explainability methods such as Grad-CAM to highlight regions of interest influencing the model’s predictions. The resulting attention maps corresponded well with regions radiologists typically inspect—such as the right lower quadrant and periappendiceal fat—indicating that the model’s reasoning process aligns with human expert interpretation. This alignment is essential for clinical validation, as interpretable AI can facilitate error analysis, improve radiologist confidence, and support educational use in medical training environments. 4.5 Limitations Despite promising results, this study has several limitations. First, the dataset size, while relatively comprehensive, remains modest for deep learning standards. Larger and more diverse datasets encompassing multicenter and multi-ethnic cohorts would help improve model robustness and external generalizability. Second, only static B-mode ultrasound images were analyzed; dynamic video sequences or cine loops might contain additional spatiotemporal cues beneficial for diagnosis. Third, although transfer learning reduced overfitting, differences in ultrasound machines, acquisition settings, and operator experience could introduce domain shifts that limit performance when applied to data from other institutions. Future research should explore domain adaptation techniques to address these issues. Moreover, while the model achieved high precision and recall, the clinical utility of false positives and false negatives must be carefully evaluated. In particular, minimizing false negatives is critical, as missed appendicitis can lead to serious complications. Integrating additional clinical and laboratory data (e.g., white blood cell count, C-reactive protein, Alvarado or PAS scores) into multimodal deep learning models may further enhance diagnostic accuracy and reduce misclassifications. 4.6 Future Directions Building upon these findings, future studies should focus on developing multimodal AI frameworks that combine ultrasound imaging with clinical metadata to emulate holistic decision-making processes. Incorporating transformer-based architectures or temporal CNNs could allow the analysis of full ultrasound video sequences rather than isolated frames, thereby capturing motion cues and probe dynamics. Additionally, explainable AI (XAI) techniques such as Layer-wise Relevance Propagation (LRP) or SHAP analysis could be used to provide quantitative interpretability, bridging the gap between AI predictions and radiological rationale. Prospective clinical trials will also be necessary to validate these systems in real-world hospital workflows and to assess how AI integration influences diagnostic speed, accuracy, and patient outcomes. 5. Conclusion In conclusion, this study demonstrates that a ResNet-based deep learning model can accurately and reliably detect pediatric appendicitis from ultrasound images, achieving strong diagnostic performance and clinical interpretability. The model’s success supports the growing evidence that deep learning can enhance pediatric imaging diagnostics, providing radiologists with advanced decision-support tools that are fast, consistent, and explainable. Continued research in data scalability, multimodal integration, and real-world deployment will be vital to fully realize the transformative potential of AI in pediatric healthcare. References Almaramhy HH (2017) Acute appendicitis in young children less than 5 years. Ital J Pediatr 43(1):15 Mostafa R, El-Atawi K (2024) Misdiagnosis of acute appendicitis cases in the emergency room. Cureus. ;16(3) Iftikhar MA, Dar SH, Rahman UA, Butt MJ, Sajjad M, Hayat U, Sultan N (2021) Comparison of Alvarado score and pediatric appendicitis score for clinical diagnosis of acute appendicitis in children—a prospective study. Annals Pediatr Surg. ;17(1) Pogorelic Z, Rak S, Mrklic I, Juric I (2015) Prospective validation of Alvarado score and Pediatric Appendicitis Score for the diagnosis of acute appendicitis in children. Pediatr Emerg Care 31(3):164–168 Mittal MK, Dayan PS, Macias CG, Bachur RG, Bennett J, Dudley NC, Bajaj L, Sinclair K, Stevenson MD, Kharbanda AB, Pediatric Emergency Medicine Collaborative Research Committee of the American Academy of Pediatrics (2013) Performance of ultrasound in the diagnosis of appendicitis in children in a multicenter cohort. Acad Emerg Med 20(7):697–702 Rahmani A, Norouzi F, Machado BL, Ghasemi F (2024) Psychiatric Neurosurgery with Advanced Imaging and Deep Brain Stimulation Techniques. Int Res Med Health Sci 7(5):63–74 Abbasi H, Afrazeh F, Ghasemi Y, Ghasemi F (2024) A shallow review of artificial intelligence applications in brain disease: stroke, Alzheimer's, and aneurysm. Int J Appl Data Sci Eng Health 1(2):32–43 Zhang C, Liu D, Huang L, Zhao Y, Chen L, Guo Y (2022) Classification of thyroid nodules by using deep learning radiomics based on ultrasound dynamic video. J Ultrasound Med 41(12):2993–3002 Norouzi F, Machado BL (2024) Predicting Mental Health Outcomes: A Machine Learning Approach to Depression, Anxiety, and Stress. Int J Appl Data Sci Eng Health 1(2):98–104 Paudyal R, Shah AD, Akin O, Do RK, Konar AS, Hatzoglou V, Mahmood U, Lee N, Wong RJ, Banerjee S, Shin J (2023) Artificial intelligence in CT and MR imaging for oncological applications. Cancers 15(9):2573 Farina JM, Pereyra M, Mahmoud AK, Scalia IG, Abbas MT, Chao CJ, Barry T, Ayoub C, Banerjee I, Arsanjani R (2023) Artificial intelligence-based prediction of cardiovascular diseases from chest radiography. J Imaging 9(11):236 Song K, Feng J, Chen D (2024) A survey on deep learning in medical ultrasound imaging. Front Phys 12:1398393 Akkus Z, Cai J, Boonrod A, Zeinoddini A, Weston AD, Philbrick KA, Erickson BJ (2019) A survey of deep-learning applications in ultrasound: Artificial intelligence–powered ultrasound for improving clinical workflow. J Am Coll Radiol 16(9):1318–1328 Park HC, Joo Y, Lee OJ, Lee K, Song TK, Choi C, Choi MH, Yoon C (2024) Automated classification of liver fibrosis stages using ultrasound imaging. BMC Med Imaging 24(1):36 Van Sloun RJ, Cohen R, Eldar YC (2019) Deep learning in ultrasound imaging. Proceedings of the IEEE. ;108(1):11–29 Marcinkevičs R, Wolfertstetter PR, Klimiene U, Chin-Cheong K, Paschke A, Zerres J, Denzinger M, Niederberger D, Wellmann S, Ozkan E, Knorr C (2024) Interpretable and intervenable ultrasonography-based machine learning models for pediatric appendicitis. Med Image Anal 91:103042 Marcinkevičs R, Reis Wolfertstetter P, Klimiene U, Chin-Cheong K, Paschke A, Zerres J, Denzinger M, Niederberger D, Wellmann S, Ozkan E, Knorr C (2023) Regensburg pediatric appendicitis dataset. (No Title). Feb 23 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7866377","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":529956817,"identity":"b735c66f-36cb-4cbb-9f35-ca74f5985718","order_by":0,"name":"Fatemeh Hosseinabadi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/0lEQVRIiWNgGAWjYBACgwNgEoIOMNgwyIH4Bx4QryWNwRisJYGgFriuNIbEBhAPr5bj7Y8/fCi4I2/OfnjjAYYEu/T5YYcfAm2xk9NtwK7F/swZM8kZBs8Md/akFQC1JOduvJ1mANSSbGx2ALsWgxs5bMw8BocZNxzIMTjA+IM5d+PsBJCWA4nbcGpJf/z5j8Fh+w3n3wD9lVCfbjg7/QMBLQkG0gwGhxM33MgBaTmcIC+dQ8AWkF96DA4n75zxrOBAQsJxww3SOUCGAR6/gELsx5/Dttv5kzd/+JBQLS8/Ox3IqLCTw6UFFSQwIEUu8UC+gRTVo2AUjIJRMBIAAHR/bk8EfeKOAAAAAElFTkSuQmCC","orcid":"","institution":"Assistant Professor of Radiology, Zahedan University of medical Sciences, Iran","correspondingAuthor":true,"prefix":"","firstName":"Fatemeh","middleName":"","lastName":"Hosseinabadi","suffix":""},{"id":529956819,"identity":"72281b3c-38ba-4d07-87d4-59bdb4709da8","order_by":1,"name":"Seyedhassan Sharifi","email":"","orcid":"","institution":"Pediatric Cardiology Subspecialist, Day General Hospital, Iran","correspondingAuthor":false,"prefix":"","firstName":"Seyedhassan","middleName":"","lastName":"Sharifi","suffix":""}],"badges":[],"createdAt":"2025-10-15 09:29:23","currentVersionCode":1,"declarations":{"humanSubjects":true,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":true,"humanSubjectConsent":true,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7866377/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7866377/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":93778217,"identity":"5d22ad0c-03fd-41d4-b315-33690b31c93e","added_by":"auto","created_at":"2025-10-17 12:46:11","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":822674,"visible":true,"origin":"","legend":"","description":"","filename":"PediatricAppendicitisDetectionfromUltrasoundImages.docx","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/f0595ef71911b4f831066f86.docx"},{"id":93778214,"identity":"734bed1c-9042-4ff2-bab5-0ca0d8a0d274","added_by":"auto","created_at":"2025-10-17 12:46:11","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs7866377.json","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/5d4148e53ad3548542c41ef6.json"},{"id":93778216,"identity":"4b80ab88-9662-4493-b963-19bb4e8408bc","added_by":"auto","created_at":"2025-10-17 12:46:11","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":59664,"visible":true,"origin":"","legend":"","description":"","filename":"rs78663770enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/2fa7695d55c95eea8f7d715f.xml"},{"id":93778220,"identity":"ef7192d1-c079-492a-92cd-3df44ca3845a","added_by":"auto","created_at":"2025-10-17 12:46:12","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":193487,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/81e1ca50e6367eb856f2b7d9.png"},{"id":93778218,"identity":"181f6bff-0b0c-4f76-b833-66f3741ba30c","added_by":"auto","created_at":"2025-10-17 12:46:11","extension":"xml","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":57107,"visible":true,"origin":"","legend":"","description":"","filename":"rs78663770structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/a98eee369d9b51fbd0dfdf98.xml"},{"id":93778219,"identity":"badf1f54-bdb8-4b86-8a88-1aa6508c8ebd","added_by":"auto","created_at":"2025-10-17 12:46:11","extension":"html","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":64266,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/ced2895d27f2e18d603b2367.html"},{"id":93779888,"identity":"076e3806-41eb-4d1d-8c1d-b9772b3a9e68","added_by":"auto","created_at":"2025-10-17 12:54:12","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":780196,"visible":true,"origin":"","legend":"\u003cp\u003eRepresentative ultrasound images from the Regensburg Pediatric Appendicitis Dataset. Top left: Ileitis showing bowel wall thickening and inflammation. Top right: Mesenterial lymphadenitis with multiple enlarged lymph nodes in the right lower quadrant. Bottom left: Appendix with surrounding tissue reaction, indicating periappendiceal inflammation and fat echogenicity. Bottom right: Appendix, visualized as a non-compressible tubular structure consistent with acute appendicitis [16,17].\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/3b44c97141f989eaac8954fd.png"},{"id":93779889,"identity":"f2ef8541-6d86-40ee-8dc5-871b9f61a5d8","added_by":"auto","created_at":"2025-10-17 12:54:17","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1508874,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7866377/v1/26e796c4-65e0-4276-8a3b-0a37f42c5760.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003ePediatric Appendicitis Detection from Ultrasound Images\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eAcute appendicitis is the most common surgical emergency in children and adolescents, accounting for approximately 1\u0026ndash;2 cases per 1,000 individuals annually. It typically results from luminal obstruction of the appendix, leading to inflammation, bacterial overgrowth, and eventual perforation if left untreated. The condition can progress rapidly from uncomplicated to complicated appendicitis, potentially resulting in peritonitis, abscess formation, and life-threatening sepsis. In pediatric patients, early and accurate diagnosis is essential because younger children often present with atypical symptoms, making clinical differentiation from other causes of abdominal pain more challenging. Studies have reported that diagnostic errors in appendicitis remain a significant concern, with misdiagnosis rates ranging from 15% to 30% in some age groups. False-negative diagnoses increase the risk of complications, while false-positive diagnoses may lead to unnecessary appendectomies, prolonged hospitalization, and higher medical costs. Consequently, improving diagnostic precision is a priority in pediatric emergency medicine [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eCurrent diagnostic workflows for suspected appendicitis combine clinical evaluation, laboratory testing, and imaging. Clinical scoring systems, such as the Alvarado Score and Pediatric Appendicitis Score (PAS), integrate symptoms (e.g., pain migration, nausea, tenderness) and laboratory markers (e.g., leukocytosis, elevated C-reactive protein) [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Although these scores are helpful for risk stratification, they are not definitive and often require confirmation through imaging studies. Among imaging modalities, ultrasound (US) is the preferred first-line technique for children due to its noninvasive nature, lack of radiation, and availability in emergency settings. However, ultrasound diagnosis of appendicitis is highly operator-dependent and subject to variability in patient anatomy, bowel gas interference, and body composition. Visualization of the appendix is successful in only 60\u0026ndash;80% of cases, and diagnostic accuracy can drop significantly in obese or uncooperative children. Even experienced radiologists may encounter difficulty differentiating early appendicitis from mesenteric lymphadenitis or gastrointestinal infections, particularly when image quality is suboptimal. These limitations create an urgent need for automated image interpretation systems capable of assisting clinicians with consistent and objective diagnostic insights [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eIn recent years, artificial intelligence (AI) and machine learning (ML) have revolutionized diagnostic imaging by allowing computational models to identify complex visual and statistical patterns beyond human perception [\u003cspan additionalcitationids=\"CR7 CR8\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. In radiology, AI has shown significant potential in tasks such as tumor detection in MRI and CT, lung pathology screening in chest X-rays, and cardiac function assessment in echocardiography [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. For ultrasound imaging specifically, AI algorithms have been successfully applied to fetal growth monitoring, thyroid nodule classification, and liver fibrosis staging [\u003cspan additionalcitationids=\"CR13 CR14\" citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. The main advantage of AI-driven systems lies in their ability to learn directly from raw image data, thereby minimizing dependence on subjective interpretations. Deep learning, particularly Convolutional Neural Networks (CNNs), has emerged as the cornerstone of medical image analysis due to its ability to automatically extract hierarchical features from low-level edges and textures to high-level structural patterns that correspond to anatomical and pathological cues. This makes CNNs particularly suitable for ultrasound data, where signal-to-noise ratios are low, and manual feature engineering is often inadequate.\u003c/p\u003e\u003cp\u003eAmong various CNN architectures, Residual Networks (ResNets) have demonstrated outstanding performance in both general computer vision and medical image analysis. ResNets introduce shortcut (skip) connections that bypass one or more layers, allowing the network to learn residual mappings instead of direct transformations. This study aims to harness the power of deep residual learning to improve the diagnostic accuracy of pediatric appendicitis detection using ultrasound data. We employed the Regensburg Pediatric Appendicitis Dataset, a comprehensive dataset collected from pediatric patients admitted with abdominal pain between 2016 and 2021 at the Children\u0026rsquo;s Hospital St. Hedwig in Regensburg, Germany. The dataset includes B-mode ultrasound images, laboratory findings, clinical scores, and expert annotations. Our primary goal was to train a ResNet-based CNN to classify ultrasound images as appendicitis or non-appendicitis and to assess its diagnostic performance against standard metrics.\u003c/p\u003e"},{"header":"2. Method","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 Dataset Description\u003c/h2\u003e\u003cp\u003eThis study utilized the Regensburg Pediatric Appendicitis Dataset, a curated clinical dataset collected retrospectively from pediatric patients admitted with abdominal pain to Children\u0026rsquo;s Hospital St. Hedwig, Regensburg, Germany, between 2016 and 2021. The dataset includes a rich combination of imaging, clinical, and laboratory information to support multimodal diagnostic modeling. Each patient record may contain one to fifteen B-mode ultrasound (US) images, captured from multiple abdominal regions of interest such as the right lower quadrant (RLQ), appendix, intestinal loops, lymph nodes, free fluid areas, and reproductive organs. The ultrasound images are stored in BMP format under the US_Pictures/ directory, with filenames corresponding to subject identifiers and view indices (e.g., 23.7.bmp for patient 23, view 7) [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eIn addition to imaging data, the accompanying file app_data.xlsx contains tabular variables summarizing laboratory test results, physical examination findings, and expert ultrasonographic assessments. Clinical scoring systems such as the Alvarado Score and Pediatric Appendicitis Score (PAS) were included to contextualize imaging results with established diagnostic criteria. Each subject is annotated for three outcome variables:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eDiagnosis \u0026ndash; appendicitis vs. no appendicitis,\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eManagement \u0026ndash; surgical vs. conservative treatment, and\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSeverity \u0026ndash; complicated vs. uncomplicated (or no appendicitis).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003e The study was approved by the Ethics Committee of the University of Regensburg (no. 18-1063-\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003e101, 18-1063_1-101 and 18-1063_2-101) and was performed following applicable guidelines and\u003c/h3\u003e\n\u003cp\u003eregulations. The ethics committee confirmed that there was no need for written informed consent\u003c/p\u003e\u003cp\u003efor the retrospective analysis and publication of anonymized routine data according to Art. 27 para.\u003c/p\u003e\n\u003ch3\u003e4 of the Bavarian Hospital Law. For patients followed up after discharge, written informed consent\u003c/h3\u003e\n\u003cp\u003e was obtained from parents or legal representatives. In this work, only the Diagnosis label (appendicitis vs. no appendicitis) was used for binary image classification.\u003c/p\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.2 Data Preprocessing\u003c/h2\u003e\u003cp\u003eThe ultrasound images exhibited high inter-patient variability in acquisition parameters such as brightness, contrast, and spatial scale. To standardize inputs, each image was converted to grayscale, resized to 224 \u0026times; 224 pixels, and normalized to zero mean and unit variance. Data augmentation was applied to improve model generalization and simulate clinical variability, including random rotations (\u0026plusmn;\u0026thinsp;10\u0026deg;), horizontal flips, contrast adjustments, and Gaussian noise injection. Because multiple images were available per subject, all images were treated as independent samples while ensuring that images from the same patient were confined to either the training or testing split to avoid data leakage. The final dataset was divided into 80% for training and 20% for testing, maintaining balanced class proportions.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e2.3 Model Architecture\u003c/h2\u003e\u003cp\u003eWe implemented a Convolutional Neural Network (CNN) based on ResNet architecture, leveraging transfer learning to benefit from pre-learned visual representations. Specifically, ResNet-50 pretrained on the ImageNet dataset was fine-tuned for the ultrasound classification task.\u003c/p\u003e\u003cp\u003eThe network structure consisted of:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eAn input layer receiving 224 \u0026times; 224 \u0026times; 3 normalized images.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eThe initial convolution and max-pooling layers from the base ResNet-50 model.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eFour residual blocks, each containing multiple convolutional layers with skip connections that enable residual learning and mitigate vanishing gradients.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eGlobal average pooling to condense spatial information.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eA fully connected dense layer with ReLU activation for feature integration.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eA final sigmoid output layer producing probabilities for the binary classes (appendicitis vs. no appendicitis).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eDuring fine-tuning, the earlier layers were frozen to preserve general low-level features, while the later residual blocks and fully connected layers were retrained to adapt to ultrasound-specific texture patterns.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e2.4 Training Procedure\u003c/h2\u003e\u003cp\u003eModel implementation and training were conducted in Python 3.10 using TensorFlow 2.14 / Keras on an NVIDIA GPU workstation. Training used the following hyperparameters:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eOptimizer: Adam (learning rate\u0026thinsp;=\u0026thinsp;1 \u0026times; 10⁻⁴)\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLoss function: Binary cross-entropy\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eBatch size: 32\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eEpochs: 100 (early stopping based on validation loss)\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eDropout: 0.3 on the dense layer to prevent overfitting\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eEach training epoch included on-the-fly data augmentation. Model checkpoints and validation metrics were recorded at each epoch. Fine-tuning the upper residual blocks improved convergence and boosted overall classification accuracy.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e2.5 Evaluation Metrics\u003c/h2\u003e\u003cp\u003eModel performance was evaluated using standard classification metrics: accuracy (ACC), precision (PRE), recall (REC), F1-score, and area under the receiver-operating characteristic curve (AUC).\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Model Performance\u003c/h2\u003e\u003cp\u003eThe proposed ResNet-based deep learning model demonstrated strong performance in detecting pediatric appendicitis from ultrasound images. After fine-tuning the pretrained ResNet-50 architecture, the model achieved an overall accuracy of 93.44%, a precision of 91.53%, and a recall (sensitivity) of 89.8% on the held-out test dataset. The F1-score, representing the balance between precision and recall, was calculated at 90.6%, indicating a stable and reliable detection capability. The Receiver Operating Characteristic (ROC) curve exhibited an Area Under the Curve (AUC) of 0.95, reflecting excellent discriminative power between appendicitis and non-appendicitis cases.\u003c/p\u003e\u003cp\u003eThe confusion matrix (Figure X) illustrated that most appendicitis cases were correctly identified, with only a small number of false negatives, primarily in borderline or low-quality ultrasound images. False positives were predominantly associated with cases showing inflamed lymph nodes or bowel wall thickening, which can mimic appendicitis sonographically. Nevertheless, the model\u0026rsquo;s high precision underscores its ability to minimize false alarms and provide radiologists with reliable assistance in triaging ambiguous cases.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Training and Validation Curves\u003c/h2\u003e\u003cp\u003eFigure X presents the training and validation accuracy and loss curves over 100 epochs. The training process exhibited steady convergence, with both training and validation accuracy improving consistently without significant overfitting. Early stopping based on validation loss prevented degradation of generalization performance. The final validation loss stabilized at 0.184, suggesting that the model effectively learned meaningful features without memorizing noise or irrelevant textures from the ultrasound data. The inclusion of dropout regularization and data augmentation contributed to stable training behavior and robust generalization.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Visual Feature Interpretation\u003c/h2\u003e\u003cp\u003eFeature activation maps generated from the final convolutional layers using Gradient-weighted Class Activation Mapping (Grad-CAM) provided qualitative insights into model interpretability (Figure X). The heatmaps revealed that the network consistently focused on anatomically relevant regions such as the appendiceal area, pericecal fat, and surrounding bowel loops, aligning well with radiologists\u0026rsquo; regions of interest during manual assessment. This correspondence between AI attention and clinical focus reinforces the physiological relevance of the learned representations, suggesting that the ResNet model\u0026rsquo;s predictions are grounded in meaningful image features rather than artifacts or background textures.\u003c/p\u003e\u003cp\u003eIn several correctly classified appendicitis cases, Grad-CAM visualizations highlighted inflamed tubular structures and peri-appendiceal fat echogenicity, both key sonographic indicators of appendiceal inflammation. In non-appendicitis cases, the model concentrated on other abdominal regions, confirming the absence of the pathological pattern. Such visualization tools enhance model transparency and can aid radiologists in understanding the reasoning behind automated classifications.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003e3.4 Comparison with Previous Approaches\u003c/h2\u003e\u003cp\u003ePrevious research on appendicitis detection using traditional machine learning methods relied primarily on handcrafted features, such as gray-level co-occurrence matrices (GLCM), edge descriptors, and statistical intensity distributions, combined with classifiers like Support Vector Machines (SVMs) or Random Forests. Reported accuracies in these methods typically ranged from 75% to 85%, limited by the subjectivity of feature engineering and the inherent variability of ultrasound image quality.\u003c/p\u003e\u003cp\u003eIn contrast, the proposed ResNet-based deep learning approach automatically extracted hierarchical spatial features directly from the ultrasound data, eliminating the need for manual feature design. This end-to-end learning framework not only improved accuracy to 93.44% but also offered superior robustness to noise, variable acquisition settings, and anatomical diversity. Furthermore, the use of transfer learning significantly reduced the amount of required labeled data and training time compared to models trained from scratch.\u003c/p\u003e\u003cp\u003eTable X summarizes the quantitative performance of the proposed model. The close alignment between accuracy, precision, and recall indicates that the model maintains balanced classification performance across both classes, avoiding bias toward either appendicitis or normal samples.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e3.5 Clinical Relevance\u003c/h2\u003e\u003cp\u003eFrom a clinical standpoint, the model\u0026rsquo;s high sensitivity (recall) is particularly valuable in reducing missed appendicitis cases, which can lead to severe complications if untreated. Likewise, the strong precision reduces the likelihood of false-positive diagnoses, which may otherwise result in unnecessary imaging or surgical intervention. Integrating such AI tools into clinical workflows could assist radiologists, especially in resource-limited or high-volume settings, by providing real-time decision support and standardized interpretation across operators.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClassification performance metric\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMetric\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eValue (%)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e93.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e91.53\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRecall (Sensitivity)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e89.80\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eF1-Score\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e90.6\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAUC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e95.0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Summary of Findings\u003c/h2\u003e\u003cp\u003eThis study developed and validated a deep residual convolutional neural network (ResNet-50) for the automatic detection of pediatric appendicitis using B-mode ultrasound images from the Regensburg Pediatric Appendicitis Dataset. The model achieved an overall accuracy of 93.44%, with a precision of 91.53% and a recall of 89.8%, demonstrating that a deep learning framework can accurately identify appendicitis in children using noninvasive imaging data. These findings highlight the potential of AI-assisted diagnostic tools to complement radiologist interpretations, particularly in emergency and resource-constrained clinical environments where rapid, objective, and reproducible results are essential.\u003c/p\u003e\u003cp\u003eThe high accuracy achieved in this study surpasses the performance reported in many traditional machine learning approaches, which often rely on handcrafted features extracted from ultrasound intensity patterns or textural statistics. By contrast, the proposed ResNet model automatically learned spatially and contextually rich representations directly from imaging data, effectively capturing the structural and morphological characteristics of the inflamed appendix and surrounding tissues.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003e4.2 Comparison with Previous Studies\u003c/h2\u003e\u003cp\u003ePrevious research efforts in automated appendicitis diagnosis have explored various imaging modalities, including CT, MRI, and ultrasound, with machine learning models such as support vector machines (SVMs), k-nearest neighbors (k-NN), and random forests. For example, studies using CT-based deep learning classifiers reported accuracies between 85% and 92%, albeit at the cost of radiation exposure \u0026mdash; a significant drawback for pediatric populations. Other ultrasound-based studies employing traditional ML approaches achieved performance typically below 85% due to the limited generalizability of manually engineered features.\u003c/p\u003e\u003cp\u003eOur findings align with recent advances in deep learning for pediatric imaging, where transfer learning using pretrained CNN architectures has shown notable improvements in diagnostic performance. The ResNet model used in this study leverages residual learning, which enables the network to train deeper architectures without the risk of gradient degradation. This design allows the model to learn both low-level ultrasound textures and high-level semantic representations critical for discriminating appendicitis from other abdominal conditions. Moreover, Grad-CAM visualization confirmed that the model\u0026rsquo;s focus regions overlapped with clinically relevant anatomical sites, lending interpretability and biological credibility to the predictions.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003e4.3 Clinical Implications\u003c/h2\u003e\u003cp\u003eAccurate diagnosis of pediatric appendicitis remains a persistent challenge, as clinical symptoms are often nonspecific and imaging results may be inconclusive. The proposed AI-driven approach has the potential to augment radiologist performance by providing a rapid, consistent, and objective assessment of ultrasound images. In emergency departments, such models could serve as second readers, flagging suspicious cases for further evaluation and helping to standardize diagnostic decisions across varying levels of clinical expertise. Importantly, this system operates entirely on noninvasive ultrasound imaging, which is safer for pediatric patients than CT-based protocols.\u003c/p\u003e\u003cp\u003eThe integration of such deep learning systems into clinical decision support platforms could reduce the diagnostic delay and variability that currently affect appendicitis management. For instance, early AI-assisted identification of appendicitis could enable faster surgical consultations, minimize unnecessary hospital admissions, and optimize the use of imaging resources. Ultimately, these tools may contribute to lowering rates of perforation and postoperative complications by facilitating timely and accurate diagnosis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003e4.4 Interpretability and Trust in AI Models\u003c/h2\u003e\u003cp\u003eOne major barrier to clinical adoption of AI systems is the lack of interpretability. Deep learning models are often viewed as \u0026ldquo;black boxes,\u0026rdquo; which can reduce clinician trust in automated outputs. To address this, the present study incorporated visual explainability methods such as Grad-CAM to highlight regions of interest influencing the model\u0026rsquo;s predictions. The resulting attention maps corresponded well with regions radiologists typically inspect\u0026mdash;such as the right lower quadrant and periappendiceal fat\u0026mdash;indicating that the model\u0026rsquo;s reasoning process aligns with human expert interpretation. This alignment is essential for clinical validation, as interpretable AI can facilitate error analysis, improve radiologist confidence, and support educational use in medical training environments.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003e4.5 Limitations\u003c/h2\u003e\u003cp\u003eDespite promising results, this study has several limitations.\u003c/p\u003e\u003cp\u003eFirst, the dataset size, while relatively comprehensive, remains modest for deep learning standards. Larger and more diverse datasets encompassing multicenter and multi-ethnic cohorts would help improve model robustness and external generalizability. Second, only static B-mode ultrasound images were analyzed; dynamic video sequences or cine loops might contain additional spatiotemporal cues beneficial for diagnosis. Third, although transfer learning reduced overfitting, differences in ultrasound machines, acquisition settings, and operator experience could introduce domain shifts that limit performance when applied to data from other institutions. Future research should explore domain adaptation techniques to address these issues.\u003c/p\u003e\u003cp\u003eMoreover, while the model achieved high precision and recall, the clinical utility of false positives and false negatives must be carefully evaluated. In particular, minimizing false negatives is critical, as missed appendicitis can lead to serious complications. Integrating additional clinical and laboratory data (e.g., white blood cell count, C-reactive protein, Alvarado or PAS scores) into multimodal deep learning models may further enhance diagnostic accuracy and reduce misclassifications.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003e4.6 Future Directions\u003c/h2\u003e\u003cp\u003eBuilding upon these findings, future studies should focus on developing multimodal AI frameworks that combine ultrasound imaging with clinical metadata to emulate holistic decision-making processes. Incorporating transformer-based architectures or temporal CNNs could allow the analysis of full ultrasound video sequences rather than isolated frames, thereby capturing motion cues and probe dynamics. Additionally, explainable AI (XAI) techniques such as Layer-wise Relevance Propagation (LRP) or SHAP analysis could be used to provide quantitative interpretability, bridging the gap between AI predictions and radiological rationale. Prospective clinical trials will also be necessary to validate these systems in real-world hospital workflows and to assess how AI integration influences diagnostic speed, accuracy, and patient outcomes.\u003c/p\u003e\u003c/div\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eIn conclusion, this study demonstrates that a ResNet-based deep learning model can accurately and reliably detect pediatric appendicitis from ultrasound images, achieving strong diagnostic performance and clinical interpretability. The model\u0026rsquo;s success supports the growing evidence that deep learning can enhance pediatric imaging diagnostics, providing radiologists with advanced decision-support tools that are fast, consistent, and explainable. Continued research in data scalability, multimodal integration, and real-world deployment will be vital to fully realize the transformative potential of AI in pediatric healthcare.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAlmaramhy HH (2017) Acute appendicitis in young children less than 5 years. Ital J Pediatr 43(1):15\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMostafa R, El-Atawi K (2024) Misdiagnosis of acute appendicitis cases in the emergency room. Cureus. ;16(3)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIftikhar MA, Dar SH, Rahman UA, Butt MJ, Sajjad M, Hayat U, Sultan N (2021) Comparison of Alvarado score and pediatric appendicitis score for clinical diagnosis of acute appendicitis in children\u0026mdash;a prospective study. Annals Pediatr Surg. ;17(1)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePogorelic Z, Rak S, Mrklic I, Juric I (2015) Prospective validation of Alvarado score and Pediatric Appendicitis Score for the diagnosis of acute appendicitis in children. Pediatr Emerg Care 31(3):164\u0026ndash;168\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMittal MK, Dayan PS, Macias CG, Bachur RG, Bennett J, Dudley NC, Bajaj L, Sinclair K, Stevenson MD, Kharbanda AB, Pediatric Emergency Medicine Collaborative Research Committee of the American Academy of Pediatrics (2013) Performance of ultrasound in the diagnosis of appendicitis in children in a multicenter cohort. Acad Emerg Med 20(7):697\u0026ndash;702\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRahmani A, Norouzi F, Machado BL, Ghasemi F (2024) Psychiatric Neurosurgery with Advanced Imaging and Deep Brain Stimulation Techniques. Int Res Med Health Sci 7(5):63\u0026ndash;74\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbbasi H, Afrazeh F, Ghasemi Y, Ghasemi F (2024) A shallow review of artificial intelligence applications in brain disease: stroke, Alzheimer's, and aneurysm. Int J Appl Data Sci Eng Health 1(2):32\u0026ndash;43\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang C, Liu D, Huang L, Zhao Y, Chen L, Guo Y (2022) Classification of thyroid nodules by using deep learning radiomics based on ultrasound dynamic video. J Ultrasound Med 41(12):2993\u0026ndash;3002\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNorouzi F, Machado BL (2024) Predicting Mental Health Outcomes: A Machine Learning Approach to Depression, Anxiety, and Stress. Int J Appl Data Sci Eng Health 1(2):98\u0026ndash;104\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePaudyal R, Shah AD, Akin O, Do RK, Konar AS, Hatzoglou V, Mahmood U, Lee N, Wong RJ, Banerjee S, Shin J (2023) Artificial intelligence in CT and MR imaging for oncological applications. Cancers 15(9):2573\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFarina JM, Pereyra M, Mahmoud AK, Scalia IG, Abbas MT, Chao CJ, Barry T, Ayoub C, Banerjee I, Arsanjani R (2023) Artificial intelligence-based prediction of cardiovascular diseases from chest radiography. J Imaging 9(11):236\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSong K, Feng J, Chen D (2024) A survey on deep learning in medical ultrasound imaging. Front Phys 12:1398393\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAkkus Z, Cai J, Boonrod A, Zeinoddini A, Weston AD, Philbrick KA, Erickson BJ (2019) A survey of deep-learning applications in ultrasound: Artificial intelligence\u0026ndash;powered ultrasound for improving clinical workflow. J Am Coll Radiol 16(9):1318\u0026ndash;1328\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePark HC, Joo Y, Lee OJ, Lee K, Song TK, Choi C, Choi MH, Yoon C (2024) Automated classification of liver fibrosis stages using ultrasound imaging. BMC Med Imaging 24(1):36\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVan Sloun RJ, Cohen R, Eldar YC (2019) Deep learning in ultrasound imaging. Proceedings of the IEEE. ;108(1):11\u0026ndash;29\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMarcinkevičs R, Wolfertstetter PR, Klimiene U, Chin-Cheong K, Paschke A, Zerres J, Denzinger M, Niederberger D, Wellmann S, Ozkan E, Knorr C (2024) Interpretable and intervenable ultrasonography-based machine learning models for pediatric appendicitis. Med Image Anal 91:103042\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMarcinkevičs R, Reis Wolfertstetter P, Klimiene U, Chin-Cheong K, Paschke A, Zerres J, Denzinger M, Niederberger D, Wellmann S, Ozkan E, Knorr C (2023) Regensburg pediatric appendicitis dataset. (No Title). Feb 23\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Zahedan University of Medical Sciences","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Pediatric Appendicitis, Ultrasound, ResNet","lastPublishedDoi":"10.21203/rs.3.rs-7866377/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7866377/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePediatric appendicitis remains one of the most common causes of acute abdominal pain in children, and its diagnosis continues to challenge clinicians due to overlapping symptoms and variable imaging quality. This study aims to develop and evaluate a deep learning model based on a pretrained ResNet architecture for automated detection of appendicitis from B-mode ultrasound images. We used the Regensburg Pediatric Appendicitis Dataset, which includes ultrasound scans, laboratory data, and clinical scores from pediatric patients admitted with abdominal pain to Children\u0026rsquo;s Hospital St. Hedwig in Regensburg, Germany (2016\u0026ndash;2021). Each subject had 1\u0026ndash;15 ultrasound views covering the right lower quadrant, appendix, lymph nodes, and related structures. For the image-based classification task, ResNet was fine-tuned to distinguish appendicitis from non-appendicitis cases. Images were preprocessed by normalization, resizing, and augmentation to enhance generalization. The proposed ResNet model achieved an overall accuracy of 93.44%, precision of 91.53%, and recall of 89.8%, demonstrating strong performance in identifying appendicitis across heterogeneous ultrasound views. The model effectively learned discriminative spatial features, overcoming challenges posed by low contrast, speckle noise, and anatomical variability in pediatric imaging.\u003c/p\u003e","manuscriptTitle":"Pediatric Appendicitis Detection from Ultrasound Images","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-17 12:46:07","doi":"10.21203/rs.3.rs-7866377/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9878dbf0-7e78-4cef-9476-41bf54557809","owner":[],"postedDate":"October 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-10-17T12:46:07+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-17 12:46:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7866377","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7866377","identity":"rs-7866377","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.