Deep Learning in Clinical Diagnostics: A Scoping Review of Innovations Shaping Future Healthcare Delivery

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Deep learning (DL) based diagnostic systems potentially offer automated image/signal interpretation and workflow support across a wide range of clinical fields, but the evidence for clinical translation is mixed. We conducted a scoping review to identify validation methods, evidence for implementation, and common methodological/operational challenges in current DL-based diagnostic research. Methods Based on PRISMA-ScR guidelines, we searched key bibliographic databases for peer-reviewed articles (2020–2025) describing DL models for diagnostic tasks producing quantitative results. Two reviewers independently screened the records and extracted the study characteristics into a standardized data extraction form (Author, year, country, domain, task, sample size, model type, comparator, metrics, validation method, setting, key findings, reported challenges). We rated each study according to its furthest advanced stage of translation (development, external validation, prospective testing, randomized trial, post-deployment). No formal risk of bias assessment was conducted. Results Twenty-four studies met the inclusion criteria across radiology (breast imaging, chest x-ray), gastroenterology (colonoscopy CADe), ophthalmology (retinal screening/prognostics), dermatology (dermoscopy), and cardiology (echocardiography/ECG). Mapping of the translation stage was found: prospective clinical testing n = 8 (33.3%), randomized trials n = 5 (20.8%), post-deployment/real-world auditing n = 4 (16.7%), external validation n = 2 (8.3%), and development/retrospective studies n = 5 (20.8%). High-performing examples from prospective or rollout studies included higher cancer detection in screening mammography, randomized evidence of higher adenoma detection with CADe, large-scale TB CXR screening at AUCs > 0.98 with workload reductions up to ~ 80%, and echocardiography automation comparable to expert metrics. Frequently cited issues were a lack of geographic/demographic diversity, spectrum bias, inconsistent external validation, underreporting of clinically relevant operating points and calibration, gaps in explainability, barriers to integration with workflow, and variable regulatory/COI transparency. Conclusion DL diagnostic platforms have attained clinical-utility evidence in many applications (screening mammography, colonoscopy CADe, programmatic CXR screening) where prospective trials or deployments are available. That said, safe widespread adoption would require standardized external validation, prospective outcome studies, and analyses of equity-focused subgroups, routine post-deployment monitoring, and transparent reporting of thresholds and governance
Full text 168,943 characters · extracted from preprint-html · click to expand
Deep Learning in Clinical Diagnostics: A Scoping Review of Innovations Shaping Future Healthcare Delivery | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Systematic Review Deep Learning in Clinical Diagnostics: A Scoping Review of Innovations Shaping Future Healthcare Delivery Godswill Uzoechina, Treasure Osajiuba This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7876598/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Background Deep learning (DL) based diagnostic systems potentially offer automated image/signal interpretation and workflow support across a wide range of clinical fields, but the evidence for clinical translation is mixed. We conducted a scoping review to identify validation methods, evidence for implementation, and common methodological/operational challenges in current DL-based diagnostic research. Methods Based on PRISMA-ScR guidelines, we searched key bibliographic databases for peer-reviewed articles (2020–2025) describing DL models for diagnostic tasks producing quantitative results. Two reviewers independently screened the records and extracted the study characteristics into a standardized data extraction form (Author, year, country, domain, task, sample size, model type, comparator, metrics, validation method, setting, key findings, reported challenges). We rated each study according to its furthest advanced stage of translation (development, external validation, prospective testing, randomized trial, post-deployment). No formal risk of bias assessment was conducted. Results Twenty-four studies met the inclusion criteria across radiology (breast imaging, chest x-ray), gastroenterology (colonoscopy CADe), ophthalmology (retinal screening/prognostics), dermatology (dermoscopy), and cardiology (echocardiography/ECG). Mapping of the translation stage was found: prospective clinical testing n = 8 (33.3%), randomized trials n = 5 (20.8%), post-deployment/real-world auditing n = 4 (16.7%), external validation n = 2 (8.3%), and development/retrospective studies n = 5 (20.8%). High-performing examples from prospective or rollout studies included higher cancer detection in screening mammography, randomized evidence of higher adenoma detection with CADe, large-scale TB CXR screening at AUCs > 0.98 with workload reductions up to ~ 80%, and echocardiography automation comparable to expert metrics. Frequently cited issues were a lack of geographic/demographic diversity, spectrum bias, inconsistent external validation, underreporting of clinically relevant operating points and calibration, gaps in explainability, barriers to integration with workflow, and variable regulatory/COI transparency. Conclusion DL diagnostic platforms have attained clinical-utility evidence in many applications (screening mammography, colonoscopy CADe, programmatic CXR screening) where prospective trials or deployments are available. That said, safe widespread adoption would require standardized external validation, prospective outcome studies, and analyses of equity-focused subgroups, routine post-deployment monitoring, and transparent reporting of thresholds and governance Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Deep learning (DL), a subset of machine learning based on representation learning through multilayer neural networks, has rapidly gained traction in clinical diagnostics, with the potential for automated image interpretation, real-time procedure guidance, signal classification, and risk-stratified screening on a large scale ( 1 , 2 , 3 ). Landmark demonstrations in dermatology and other image-rich domains have provided evidence that DL can achieve expert-level performance on specialized tasks, while large-scale endeavors focused on applying DL to electronic health records and population screening demonstrate applicability across modalities and data types ( 2 , 3 ). Although substantial accuracies are reported in many single-study reports, the translational pathway from algorithmic development to safe, effective clinical implementation has been uneven. Systematic reviews describe significant heterogeneity in study designs, lack of consistent external validation, and pervasive differences in reporting standards; factors that can inflate apparent performance and limit confidence in real-world applicability ( 4 , 5 ). The divide between retrospective development and operational deployment is further widened by datasets limitations that allow for hidden stratification and other modes of failure: models that have strong average performance may still fail systematically on clinically relevant but under-represented subgroups or imaging subsets ( 6 ). Bias, fairness, and robustness concerns have led to renewed focus on dataset auditing, subgroup reporting, and other formal procedures for identifying and addressing distributional and label-quality problems ( 7 ). Regulators and governing bodies have also hinted towards the requirement for lifecycle methodologies for AI/ML medical software, emphasising pre-specified change-control plans, ongoing performance monitoring, and transparent evidence of clinical safety as necessary conditions for deployment ( 8 ). At the same time, reporting guidelines for diagnostic AI are being developed to define a standard way to report accuracy, thresholds, calibration, and implementation settings to allow clinicians, payors, and regulators to better assess readiness for use ( 9 ). Collectively, these changes thus present an urgent demand for mapping the current evidence landscape: which clinical fields have advanced beyond retrospective development to prospective evaluation or real-world implementation, what range of potential outcomes beyond diagnostic accuracy (such as workflow impact, safety indicators, or clinician-assistance advantages) have been assessed, and which cross-cutting methodological and implementation challenges recur across studies ( 10 ). The scoping review aims to systematically map published DL diagnostic studies, classify their validation and implementation phase, and integrate operational and equity findings to guide researchers, clinicians, and policy makers on where DL is ready for clinical support and where further work remains. Methodology Study Design We performed a scoping review to map the current literature on deep-learning (DL) based diagnostic systems and assess their clinical validation and implementation. The review was conducted in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews) statement to ensure transparent and reproducible reporting of objectives, search methodology, selection, and data extraction ( 11 ). The review protocol, which outlined the inclusion criteria, search strategy, and data-charting elements, was established in advance. Eligibility criteria We considered primary empirical studies published in peer-reviewed journals, which described the development, validation, or clinical application of a deep learning based model for a diagnostic task in human healthcare. Eligible study designs included retrospective development study, prospective evaluation, randomized controlled trial, quasi-randomized trial, post-deployment audit, or real-world application report. We considered studies across modalities (imaging, signals, endoscopy, dermoscopy, ECG, etc.) so long as the studies reported at least one quantitative diagnostic or implementation outcome (e.g., AUC, sensitivity, specificity, detection rate, ADR, workflow/time metrics). Excluded were review articles, editorials, purely technical papers with no clinical data, animal research, and studies applying only classical (non-deep) machine learning methods. We restricted ourselves to English-language full texts. The eligible time window was 2022–2025 to cover recent DL methodologies that could be translated into clinical applications. Information sources and search strategy We conducted a search of PubMed and Google Scholar for articles published between 1 January 2020 and 30 September 2025. The search combined MeSH terms and free text terms relating to three concepts: ( 1 ) deep learning/convolutional neural networks/transformer; ( 2 ) diagnosis/screening/detection/classification; and ( 3 ) validation/implementation/trial/deployment. An illustration search string for PubMed was: ("deep learning" OR "convolutional neural network" OR "CNN" OR "transformer" OR "neural network") AND ("diagnosis" OR "screening" OR "detection" OR "classification") AND ("validation" OR "prospective" OR "randomized" OR "trial" OR "deployment" OR "implementation"). Secondary search methods included manual handsearching, including backward and forward citation tracking of eligible studies, and screening the reference lists of relevant reviews Study selection All results obtained were loaded into a citation manager, and duplicates were removed. Title/abstract screening was conducted by two independent reviewers using pre-defined inclusion criteria; potentially relevant full texts were subsequently retrieved for a duplicate assessment. Differences in opinion at any screening level were resolved by consensus; further disagreements were mediated by a third reviewer. We recorded the full-text exclusions with reasons for exclusion, and the selection flow is presented as a PRISMA-ScR flow diagram (Fig. 1). Data charting We created a structured data spreadsheet in Microsoft Excel for data extraction and charting, which corresponded to the columns of the provided extraction table. From each study, we extracted the following details verbatim where possible: Author(s) and year of publication; DOI/URL if available; Country; Clinical domain; Diagnostic task/modality; Sample size (N) and dataset source; Deep learning model type (architecture/ensemble/commercial vendor name); Comparator (human readers, standard of care, other algorithms); Outcome metrics reported (AUC, sensitivity/specificity, detection rates, DSC, MAE, reading time, workload metrics, etc.). Validation approach (internal train/validation/test, external test, temporal split, prospective test, randomized design, post-deployment audit); Setting of Application (single centre, multicenter, screening program, deployment center); Main Findings (primary interpretations of authors); Reported Challenges/Limitations. One reviewer conducted the initial extraction, and a second reviewer independently verified all extracted fields for completeness and accuracy. Discrepancies were resolved by discussion and, when necessary, review of the full text. Figure 1. PRISMA-ScR flow diagram summarising study selection. From an initial 2,134 records, 24 studies met the inclusion criteria and were included in the final scoping review. Figure 1: PRISMA-ScR flowchart Results Study selection We identified the body of literature that met our pre-specified inclusion criteria through both database and grey-literature searches (see Methods). Following the removal of duplicates and screening of titles/abstracts, full texts were assessed for eligibility, and findings were mapped for studies that reported a deep-learning method applied to a human clinical diagnostic task. The study selection process is summarized in the PRISMA-ScR flow diagram (see Methodology) Study characteristics The included studies covered various clinical domains, with study designs ranging from retrospective development studies to large prospective deployments and randomized trials. Table 1 reports a summary of the study domains, sample size, dataset sources, deep learning architectures, validation methods, and the stage of clinical translation. Table 1: Study characteristics Clinical domains most frequently highlighted were radiology (particularly breast imaging and chest radiography), gastroenterology (colonoscopy polyp detection), ophthalmology (fundus-based screening and prognostics), dermatology (dermoscopy), and cardiology (echocardiography and ECG). Pulmonology (TB detection on CXR) was also emerging as a field with large datasets and post-deployment evaluations. The size of the datasets ranged from small clinical studies (hundreds of patients/images) to very large screening datasets (hundreds of thousands to over one million images in deployment settings). Model types covered traditional CNN variants (ResNet, Xception, EfficientNet), ensemble techniques, multi-task models with temporal transformers (echocardiography), object-detection architectures (RetinaNet), as well as explainable/gradient-based visualizations (Grad-CAM). A few of the studies reported commercially available systems (e.g., Lunit, ScreenTrustCAD, GI Genius, ARDA), while others used open-source or custom models. Validation methodologies varied from retrospective training/testing splits to multi-center external validation, prospective trials, randomized controlled trials, and extensive post-deployment audit studies. Table 2 summarizes studies by clinical domains reporting counts, common diagnostic tasks, median/mean sample size, and median primary performance metric (e.g., AUC or ADR where applicable) for each clinical domain. Table 2: Summary by clinical domain. Clinical domain Number of included studies (n) Common diagnostic tasks (typical modalities) Median sample size (per study) — range (min–max) Common primary performance metric(s) reported Median performance (metric + value) — basis (n studies used) Notes/interpretation Radiology: Breast imaging 6 Screening mammography (AI as additional/independent reader); CEM lesion segmentation/classification; AI triage to MRI 40,062 (range 1,912 – 463,094) Cancer detection rate (CDR) / relative % change; AUC for lesion classification (when reported) Median relative increase in cancer detection ≈ 13.8% (computed across studies that reported relative or % change: Eisemann +17.6%, Chang +13.8%, Dembrower +4% → median 13.8%, n=3). Single-study AUC example: Zheng et al. AUC ≈0.94 (internal/external). Strong translation activity: several large prospective/rollout studies and RCTs; heterogeneous metrics (detection rates, PPV, recall) so use domain-specific metrics rather than pooled AUC. Gastroenterology: Colonoscopy (CADe) 4 Real-time polyp/adenoma detection (video/white-light) 735 (range 223 – 805) Adenoma detection rate (ADR), Polyp detection rate (PDR), Adenomas per colonoscopy (APC), miss rates Median ADR with AI ≈ 43.7% (AI ADRs: 50.44%, 37%, 35%, 59.1% → median 43.72%, n=4). Median control ADR ≈ 35.8% (controls median 35.82%, n=4). Median absolute ADR improvement ≈ 7.9 percentage points (n=4). Consistent RCT evidence of clinically meaningful ADR/PDR improvements across multiple centers; median ADR improvement ~8 pp. FP alerts and withdrawal time reported variably. Dermatology: Dermoscopy / Melanoma detection 3 Dermoscopy-based melanoma vs benign lesion classification; clinician + AI assistance 435 patients (range 228 lesions – 13,900 images) AUC; sensitivity / specificity; accuracy Median AUC ≈ 0.945 (AUC values used: 0.816; 0.9211; 0.968; 0.9868 — median ≈ 0.945, n≈4 AUC values from 3 studies). Typical reported sensitivity (with AI) often high (≈95–97%) (n studies vary). High sensitivity in many studies but specificity variable (trade-offs). Dataset diversity (skin tones) limited in several studies and many results are dataset/computational rather than large prospective deployments. Ophthalmology: Fundus screening & prognostics 4 Multi-disease fundus screening; diabetic retinopathy screening; prognostic (time-to-progression) models 18,760 (range 4,537 – 110,784) — per-study sizes (individuals/images) Screening: sensitivity/specificity; Prognostics: C-index (survival) Median screening sensitivity ≈ 91.4% (screening studies: Dong 89.8%, Brant 97.0%, Paisan 91.4% → median 91.4%, n=3). Prognostic (DeepDR Plus) median C-index ≈ 0.84 (external range 0.823–0.862; use midpoint ≈ 0.84, n=1 study with multiple cohorts). Mix of massive population screening deployments (high external validity) and prognostic models enabling personalized screening intervals. Strong external validations but geographic training concentration (China/India) noted as a limitation. Pulmonology / Chest radiography: TB & general CXR 3 TB detection on chest X-ray; segmentation/localization (Grad-CAM) 12,848 (range ~2,000 – 1,040,000) AUC; accuracy; false-negative rate (FNR) / workload reduction (WLR) Median AUC ≈ 0.992 (AUCs: Sharma 0.999; Munjal 0.9851 → median ≈ 0.992, n=2 AUCs). Other studies report extremely high accuracy/accuracy ~99–100% (Mujeeb). Some studies are very large real-world deployments (Munjal ~1.04M CXRs) demonstrating operational safety and large workload reductions; many high AUC/accuracy numbers come from public datasets or retrospective analyses — prospective generalizability needs ongoing evaluation. Cardiology: Echocardiography (TTE) 2 Automated full TTE interpretation (view classification, quantification) 48,146 (median of reported per-study values: 32,265 and 64,028 → median 48,146; range 32,265 – 64,028) AUC for classification tasks; MAE for quantitative estimates (LVEF MAE) Median classification AUC ≈ 0.94 (Holste median AUC 0.91; Emily S. view AUROC >0.97 → median ≈ 0.94, n=2 studies). LVEF MAE ≈ 4.2–4.5% reported in examples. Very large video datasets and multi-task models with external validations; results show near-expert performance for many labels and good quantification accuracy, but prospective outcome studies and workflow integration remain limited. Cardiology: ECG arrhythmia classification 2 12-lead ECG automated arrhythmia detection / multi-class classification 10,943 (median of 48 (MIT-BIH) and 21,837 (PTB-XL) → 10,943; range 48 – 21,837) Accuracy; AUC (when reported) Median accuracy / AUC ≈ high (≈ 98%) — examples: Bai (MIT-BIH / PTB generalization) accuracy/AUC ~99.4% / 0.9875; Atwa (PTB-XL) multiclass accuracy typically 71–98% depending on task — median across available figures ≈ ~98% (binary/superclass) (n=2 studies; metric differs by classification granularity). Benchmark datasets show near-perfect classification on some tasks, but MIT-BIH is small and older — real-world generalizability requires larger clinical dataset validation and prospective testing. Figure 2 highlights annual counts of included studies, stacked by clinical domain, demonstrating the temporal growth and domain shifts in DL diagnostic research. Figure 2: Timeline of publications by domain Thematic synthesis by clinical domain Radiology: Breast imaging and mammography screening A significant portion of the included studies assessed DL for breast imaging, with works on automated lesion segmentation, single-read support, AI as independent readers in population screening, and AI triage for supplemental MRI. Representative findings include: Segmentation & single-mass classification (contrast-enhanced mammography): Zheng et al., 2023 proposed and evaluated a Fully Automated Pipeline System (RefineNet + Xception + Pyramid Pooling) for a cohort of 1,912 women and achieved Dice similarity coefficients of 0.888 (internal), 0.820 (external), prospective 0.837, and AUC of 0.947 (internal) and 0.940 (external); the system decreased reading time (6 seconds per case vs ~3 minutes) and enhanced radiologist performance and BI-RADS reclassification in ~12–13% of the cases. The study demonstrates significant segmentation/classification accuracy as well as practical workflow benefits (12). AI as an additional or replacement reader in population screening: Multiple large prospective and real-world studies highlight that AI can safely augment or replace one human reader: Ng A.Y. et al., 2023 (Hungary) stated that the integration of an ensemble commercial AI ("Mia") led to an increase in cancer detection by 0.7–1.6 per 1,000 and a slight recall increase (0.16–0.30%). The study matured from pilot to live rollout and demonstrated feasibility in a screening program (13). Dembrower et al., 2025 (Sweden) screened 55,581 women and demonstrated that AI alone or displacing one radiologist with AI was non-inferior in cancer detection rate compared to two-radiologist reads; triple reading (2 radiologists + AI) increased cancer detection by 8% (14). Eisemann et al, 2025 (Germany), AI-assisted double reading significantly increased cancer detection (6.7–5.7 per 1,000; +17.6%), and biopsy PPV also improved, across 463,094 screens performed (15). AI triage for MRI: Salim et al., 2024 applied AI to triage the top ~6.9% for adjunct MRI, detecting 64.4 cancers/1,000 MRIs as opposed to 16.5/1,000 with density-based triage; a nearly 4× increase in yield per MRI (16). Many of the breast imaging studies are prospective, population-based, or linked to national programs and are among the highest levels of translation in our corpus (see Table 3 and Figure 4). A common limitation is the restricted geographic variability (many are from Sweden, South Korea, Hungary, and China) and reliance on a few commercial systems. Table 3 reports counts and proportions of studies in each translation stage: development only, internal validation, external validation, prospective clinical testing, randomized trials, and post-deployment surveillance/real-world audits. Table 3: Deployment readiness across included studies Translation stage (highest achieved) Count (n) % of studies Example studies (representative; author, year) Prospective clinical testing (prospective evaluations, paired-reader, prospective cohort studies, prospective external validation) 8 33.3% Zheng T et al., 2023; Dong L et al., 2022; Dembrower K et al., 2025; Chang Y-W et al., 2025; Marchetti MA et al., 2023; Winkler JK et al., 2023; Yifan Peng et al., 2024; Paisan Ruamviboonsuk et al., 2022 Randomized trials (RCTs and quasi-randomized trials) 5 20.8% Glissen Brown et al., 2022 (RCT); Park D.K. et al., 2024 (RCT); Latcher et al., 2023 (RCT); Salim M. et al., 2024 (RCT); Lagström RMB et al., 2025 (quasi-randomized) Post-deployment surveillance / real-world audits / large rollouts 4 16.7% Ng A.Y. et al., 2023 (rollout); Munjal P. et al., 2025 (large screening deployment); Brant A. et al., 2025 (post-deployment adjudication); Eisemann N. et al., 2025 (multi-site real-world implementation) External validation (multi-center or multi-cohort external testing, but not prospective clinical testing) 2 8.3% Holste G. et al., 2025; Emily S., 2023 Development only / computational (retrospective train/val/test, public-dataset benchmarks) 5 20.8% Mahmud M.A.A., 2025; Sharma V. et al., 2024; Mujeeb et al., 2024; Xiangyun Bai, 2024; Ahmed E.M. Atwa et al., 2025 Total 24 100% — Gastroenterology: Colonoscopy polyp/adenoma detection (CADe) Multiple randomized and quasi-randomized trials reported clinically meaningful rises in adenoma detection rate (ADR), reductions in miss rates, and increases in polyps detected per colonoscopy (APC). Representative findings are: Randomized tandem or parallel RCTs: Glissen Brown et al., 2022 (US, EndoScreener, SegNet) in 223 patients decreased adenoma miss rate (20.12% vs 31.25%, P=0.0247) as well as polyp miss rate (20.7% vs 33.7%). First-pass ADR was trended higher with CADe (17). Latcher et al., 2023 (Israel, DEEP² CADe) reported ADR 37% vs 27% (+10%, p=0.0057) and PDR 55.5% vs 38.7% (p<0.001) (18). Park D.K. et al., 2024 (Korea, RetinaNet) revealed PDR 62% vs 52% (p=0.01) and ADR 35% vs 28% (p=0.03) (19). Large quasi-randomized trial / real-world rollout: Lagström RMB et al., 2025 (Denmark, GI Genius v3) in 795 patients highlighted ADR 59.1% vs 46.6% (p<0.001) (20). CADe systems usually operate in real-time, considering false-positive (FP) alert rates and withdrawal time. Many studies report a modest rise in procedure time or no change; FP alerts per exam are sometimes quantified (e.g., ~4 false alerts/exam - Latcher et al.). The body of randomized evidence ranks colonoscopic CADe among the most solidly validated DL diagnostic applications. Ophthalmology: Fundus-image screening and prognostics Ophthalmology studies included multi-disease classifiers for screening and DL prognostic models that predict time-to-progression. Finings are: Multi-disease screening (RAIDS): Dong et al., 2022 trained RAIDS on 120,002 development images and prospectively validated its performance on 208,758 images (110,784 individuals). RAIDS sensitivity for any of the 10 retinal diseases was 89.8% with per-disease accuracy ranging up to 95–99.9%; RAIDS was comparable or better than human graders for several diseases in multicenter screening practice (21). Post-deployment audit (ARDA): Brant et al., 2025 assessed ~4,537 adjudicated images sampled from ~600,000 screened images from the Aravind program and demonstrated severe NPDR/PDR sensitivity at 97.0% and specificity at 96.4%, indicating robust post-deployment performance for a CE-approved ensemble (22). Prognostics: DeepDR Plus (Yifan Peng et al., 2024), a model trained on 717,308 images, externally validated on 8 longitudinal cohorts, reporting C-indices for 1–5 year DR progression prediction of ~0.82–0.86; a prospective real-world validation demonstrated safely prolonged screening intervals with minimal delayed detection (23). Ophthalmology includes both large prospective screening deployments and prognostic models that have been externally validated. Authors’ concerns relate to generalizability to areas beyond the regions where they were trained (mainly China/India) and reliance on image quality. Dermatology: Dermoscopy and melanoma detection A few prospective or diagnostic results on CNN assistance for dermoscopic melanoma classification were obtained in dermatology: they revealed high sensitivity, but only moderate specificity. Findings were: Marchetti et al., 2023 (ADAE): prospective observational study (603 biopsied lesions, 95 melanomas) in which ADAE obtained a sensitivity of 96.8% and increased dermatologist AUC from 0.780 → 0.816 (p=0.042); however, specificity was still low (37.4%) (24). Winkler et al., 2023 (Moleanalyzer Pro): sensitivity of the dermatologist was increased from 84.2% to 100% by CNN assistance in 228 lesions; specificity and accuracy also improved (25). Mahmud et al., 2025 reported high accuracy and AUC across large public dermoscopy datasets using an Xception-based model with explainability, but this was computational (no clinical deployment) and raised concerns about skin-tone representation (26). Cardiology: Echocardiography and ECG Studies in cardiology ranged from high-volume automated transthoracic echocardiogram (TTE) interpretation to ECG arrhythmia detection. Representative findings were: Echocardiography (PanEcho): Trost et al., 2025 trained a model on ~1.2 million echo videos and reported a median classification AUC of 0.91 (IQR 0.88–0.93) and LVEF MAE ≈4.2–4.5% (internal/external), indicating multi-task automation across numerous labels and extension to abbreviated or POC echo (27). Automated quantification and prognostic models: Emily S., 2023 (DROID workflow) reported AUROCs >0.97 for view classification and showed prognostic associations (e.g., HR 1.43 per 1 SD decrease in model-derived LVEF for heart failure) (28). ECG arrhythmia classification: Atwa et al., 2025 and Xiangyun Bai, 2024 achieved very high accuracies in benchmark datasets (MIT-BIH, PTB, PTB-XL), with hybrid CNN + attention approaches reporting >99% in certain binary cases and promising multiclass results (29,30). Pulmonology: Chest X-ray (TB screening and broader chest CXR) Pulmonology studies included large-scale screening applications as well as cross-dataset validations for TB identification. Findings were: Large-scale deployment (AIRIS-TB): Munjal et al., 2025 evaluated AIRIS-TB on ~1.04 million CXRs with an AUC of 98.51% and TB-FNR of 0% (safe setting), reporting up to 80% automated reporting capability without compromising safety (31). Other high-volume screening: Sharma et al., 2024, and Mujeeb et al., 2024 reported very high accuracies for TB detection using UNet segmentation + Xception classification and lightweight CNNs, respectively; however, they were constrained to public datasets or smaller curated sets (32,33). Cross-cutting themes and methodological observations First, with respect to the rigor of validation and readiness for translation, several clinical domains, most notably breast cancer screening, colonoscopy computer-aided detection (CADe), and some ophthalmology screening programs, demonstrate mature levels of clinical validation. These include rigorously designed randomized controlled trials (RCTs) and even large-scale national rollouts. Notable examples include Dembrower et al. (2025), Eisemann et al. (2025), Lagström et al. (2025), and Ng et al. (2023), which demonstrated that deep learning (DL) systems can be successfully implemented in real-life clinical environments. But many studies remain confined to the retrospective computational phase, relying primarily on internal or public datasets with conventional train-validation-test splitting and lacking external validation or clinical evaluation. Typical examples are Mahmud (2025) in dermoscopy and Mujeeb et al. (2024) in tuberculosis chest X-ray (CXR) classification. This is visualised in Figure 3, which shows the proportion of studies that attained external validation, prospective testing, RCT evidence, or post-deployment evaluation, and also reveals differences in translation readiness across domains. Figure 3: Illustration of the proportion of studies achieving external validation, prospective testing, RCT evidence, or post-deployment evaluation Concerning explainability and interpretability practices, the majority of the studies featured qualitative explainability tools such as Grad-CAM, saliency maps, or confidence overlay, especially in dermoscopy, chest radiography, and TB imaging (Mahmud 2025; Sharma 2024; AIRIS-TB). Nevertheless, these visualizations were often expressed in a purely descriptive manner without structured human-factors evaluation or quantitative assessment of interpretability. In studies where explainability techniques were tightly coupled with clinical workflows (e.g., Grad-CAM-based overlays utilised by radiologists or dermatologists), authors noted that such tools facilitated clinician confidence and transparency, without supplanting human adjudication or expert oversight. A persistent deficiency was identified in terms of dataset diversity, fairness, and demographic reporting (Figure 4 highlights the number of included studies by country). Many studies neglected to provide detailed reports on the demographic composition of the datasets, including race, ethnicity, skin tone, or socio-economic status, or on the performance metrics of specific subgroups. This limitation was significantly evident in dermatology and some radiology studies. In the relatively few cases where subgroup analyses were conducted, performance often changed substantially between demographic subgroups, highlighting the limited generalizability of the model and the need for diverse, representative external validation. Figure 4: Number of Included Studies by Country Concerning reporting quality and reproducibility, only a handful of studies rendered their source code or model weights publicly available. A lot of them relied on proprietary algorithms or commercial platforms, which hindered independent validation and objective benchmarking. Reporting practices were similarly inconsistent: while AUC (Area Under the Curve) was almost universally reported, some details, such as decision thresholds, calibration methods, and measures of uncertainty, were frequently unreported. Sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) at clinically relevant thresholds were rarely reported. Implementation outcomes & reported challenges Among the included studies, implementation-related outcomes mainly focused on the efficiency of operation, diagnostic performance, and safety, representing both the potential and challenges of real-world AI application in clinical practice. The operational outcomes reported indicated a considerable promise in accelerating workflows and redistributing workloads. For example, Zheng et al. (2023) estimated that the fully automated FAPS classification took about six seconds per case compared to nearly three minutes of radiologist review, illustrating significant time savings. In the same manner, AIRIS-TB deployment as reported by Munjal et al. (2025) achieved up to 80% automated reporting with sustained diagnostic safety. However, not all findings reflected workload reduction: Ng et al. (2023) noted a 4–6% workload increase when implementing AI as a second reader in screening workflows, and Eisemann et al. (2025) found slight increases in reading volume along with enhanced cancer detection rates and positive predictive value. Figure 5 illustrates the distribution of study counts across deep learning model families and clinical domains, showing both the frequency of utilization and corresponding mean diagnostic performance (AUC) within each category. Figure 5. Distribution of Deep Learning Model Families Across Clinical Domains with Representative AUC Values Diagnostic and management-related impacts were also described. Zheng et al. (2023) reported that AI-assisted mammographic interpretation led to BI-RADS category changes that altered management decisions in about 12–13% of patients, highlighting AI’s real impact on downstream clinical care. In another case, Salim et al. (2024) reported a 4-fold higher diagnostic yield per MRI scan using AI-based MRI triage. In addition to detection, prognostic tools, for instance DeepDR Plus (Yifan Peng et al., 2024), extended the scope of AI to outcome forecasting; predicting diabetic retinopathy progression risk and safely extended screening intervals with a ≤0.18% risk rate of missing vision-threatening retinopathy at detection to a whole new level of outcome forecasting—predicting diabetic retinopathy progression risk and safely extending screening intervals with a ≤0.18% risk rate of delayed vision-threatening retinopathy detection. Taken together, these examples show that well-calibrated models can not only improve diagnostic efficiency but also optimize resource utilization and personalize screening strategies. Figure 6 presents a Sankey diagram highlighting how dataset sources link to different deep learning model families and their corresponding validation stages, illustrating the progression of data utilization from public or clinical origins to various levels of model evaluation and implementation. Figure 6. Flow of Dataset Sources Through Model Families to Validation Stages Despite these encouraging results, several implementation challenges and barriers were invariably reported. Generalizability and external validity were flagged as significant concerns: many models had lower performance when tested in populations that differed geographically, ethnically, or in imaging modality. Many authors highlighted the importance of diverse, multi-ethnic external validation cohorts. Concerns about dataset and spectrum bias were similarly raised, stating that publicly available datasets or datasets curated by research groups may not truly capture the real-world heterogeneity of disease, leading to the inflation of performance estimates. Explainability and clinician acceptance posed further barriers. While visualization methods (such as saliency or Grad-CAM maps) have improved the interpretability of the models, they are not sufficient on their own to establish full clinical trust, prompting suggestions for structured human-factors studies and workflow-integrated user interfaces. Infrastructural and operational difficulties were reported as well: how best to integrate PACS with EHR systems, computational latency for real-time analysis, and the challenge of instituting escalation pathways for AI-flagged cases were persistent issues in many of the reports. Finally, the considerations for regulation and ethics were uneven. Many studies involved commercial algorithms or industry collaboration, but inconsistently reported conflicts of interest, and regulatory status such as CE marking, FDA clearance. Resource-related constraints were especially relevant in low- and middle-income environments, where the cost of hardware and connectivity played a role in shaping the decisions around model design. Several teams explicitly focused on lightweight or edge-compatible models, for example, Mujeeb et al. (2024), who optimized for constrained computational infrastructure, to enable equitable. Discussions Principal Findings This scoping review demonstrated that current deep-learning (DL) diagnostic research varies widely in both maturity and the quality of evidence. Certain domains, such as screening mammography, real-time colonoscopy CADe, and large-scale chest X-ray tuberculosis (TB) screening, have moved on from single-centre development to prospective evaluations, randomized trials, and national or programmatic rollouts that consistently demonstrate positive clinical impact on detection rates, workflow throughput, or both ( 12 – 15 , 17 , 24 , 31 ). At the same time, a large number of studies are still retrospective or computational (train/validation/test on public or curated datasets), boasting high performance metrics that have not been externally validated or assessed in operational settings ( 14 , 19 , 21 ). In general, while several DL systems have been shown to produce compelling diagnostic performance in controlled and some real-world implementations, the evidence base is variable across modalities and geographies. Interpretation: innovation maturity across domains. The collective findings indicate a spectrum of maturity. Breast imaging has the most mature programmatic evidence: there are several large, prospective, population-based or real-world demonstration studies indicating that artificial intelligence (AI) can augment or replace a reader without loss of cancer detection and with potential gains in positive predictive value and efficiency ( 13 – 15 ). Gastroenterology CADe for polyp/adenoma detection is supported by randomized and multicenter trials indicating consistent improvements in adenoma and polyp detection rate, and a corresponding decrease in miss rates ( 17 , 19 , 20 ). A multitude of large-scale CXR TB screening solutions have matured to operational large-scale deployments, reporting high AUCs with measurable reductions in workload at an adapted operating point in screening workflows ( 31 ). In contrast, many dermatology, single-task radiology, ECG, and some prognostic applications are still largely in the development/retrospective validation stage: strong internal results are common, but independent external validation, prospective implementation studies, and long-term outcomes data are scarce or non-existent ( 21 , 22 , 25 ). Hence, the level of maturity of innovation is closely correlated with (a) the availability of large, programmatic datasets and (b) the existence of pragmatic prospective evaluations. Clinical implications On current evidence, DL approaches are poised to enhance and support clinical services in specific, constrained tasks for which prospective and rollout data are available. These are ( 1 ) screening mammography workflows: as an independent reader, as a safety net, or triage in organized screening programs to substantially increase and faclitate cancer detection while reducing human workload under controlled deployment protocols ( 13 – 15 ); ( 2 ) population or programmatic CXR TB screening where high-throughput implementations have demonstrated safe triage thresholds and very high sensitivity with the potential for substantial automation ( 31 ); and ( 3 ) CADe for GI endoscopy for which RCT evidence supports routine use to reduce adenoma miss rates and enhance adenoma detection measures in multiple centres ( 17 , 19 , 20 ). For other tasks (dermoscopy, single-centre radiology algorithms, ECG arrhythmia classifiers, echocardiography prognostic models), DL may be suitable for decision-support or augmentation in research-savvy or pilot clinical environments but should not be considered for widespread unguided replacement of clinician judgment until comprehensive external validations, prospective effectiveness studies, and implementation activities have been conducted ( 21 , 22 , 24 , 25 ). Research and policy recommendations. To rapidly facilitate safe and equitable translation, the field should give priority to: ( 1 ) multicenter, pragmatic randomized trials and prospective implementation studies that assess patient-relevant and cost-effectiveness outcomes rather than algorithmic metric performance alone; ( 2 ) standardized validation hierarchies that mandate independent external validation cohorts, temporal validation, and the reporting of calibration and clinically relevant operating points; ( 3 ) prespecified subgroup analyses by key demographic and technical covariates (age, sex, race/ethnicity, device/vendor) and public dissemination of such findings; ( 4 ) standardized postdeployment surveillance pipelines (continuous performance monitoring, drift detection, threshold recalibration) with publicly accessible run logs of model updates; ( 5 ) human-factors research to assess the impact of explainability tools (eg, saliency maps) on clinician decision-making and workflow; and ( 6 ) well-defined regulatory pathways and conflict-of-interest disclosure policies that facilitate independent audit of commercial systems in clinical environments. Where feasible, funders and journals should mandate data sharing or, at a minimum, standardized metadata descriptors to allow for replication and pooled analyses. Strengths and limitations of this scoping review. The strengths of this review include the detailed extraction of study design, validation method, implementation environment, and clinical outcome from 24 recent DL diagnostic studies, with an emphasis on the translation stage and real-world implementation. Additionally, it integrates common methodological and ethical issues that have a direct bearing on decision-making regarding implementation. Limitations are the scoping design (a quantitative meta-analysis was not conducted), potential publication and language bias, inclusion of diverse types of studies (clinical trials, prospective audits, computational studies) that impair comparability, and dependence on reporting within primary studies (which occasionally lacked complete metrics disclosures or subgroup data). Finally, there is some indication that new deployment studies may shift the evidence balance, given the rapidly evolving nature of the literature; ongoing updates and living evidence syntheses will be critical . Conclusion Deep learning has been shown to reach clinically useful thresholds for selected diagnostic scenarios, including routine screening mammography, programmatic CXR TB screening, and colonoscopy CADe; in these areas, pragmatic large-scale evidence supports cautious adoption with operational safeguards. That said, for several other interesting applications, far more methodological bolstering, more extensive external validation, prospective outcome trials, equity-centered subgroup analyses, and mandated post-deployment surveillance are warranted before recommending widespread, unsupervised clinical adoption. Declarations Ethics approval and consent to participate Not applicable. This study did not involve human participants, human data, or tissue. Consent for publication Not applicable. No person’s data in any form is presented in this manuscript. Competing interests All authors certify that they have no affiliations with or involvement in any organization or entity with any financial interest (such as honoraria, educational grants, participation in speakers’ bureaus, membership, employment, consultancies, stock ownership, or other equity interest; and expert testimony or patent-licensing arrangements), or non-financial interest (such as personal or professional relationships, affiliations, knowledge or beliefs) in the subject matter or materials discussed in this manuscript. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Author Contribution GU conceived the study, designed the methodology, conducted the literature search, performed data extraction and synthesis, and drafted the manuscript. TO contributed to data verification, interpretation of findings, and drafting and critical revision of the final manuscript. Both authors reviewed and approved the final version for submission. GU is the guarantor of the manuscript. References Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44–56. 10.1038/s41591-018-0300-7 . Epub 2019 Jan 7. PMID: 30617339. Esteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115–118. 10.1038/nature21056 . Epub 2017 Jan 25. Erratum in: Nature. 2017;546(7660):686. doi: 10.1038/nature22985. PMID: 28117445; PMCID: PMC8382232. Rajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1:18. 10.1038/s41746-018-0029-1 . Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44–56. 10.1038/s41591-018-0300-7 . Epub 2019 Jan 7. PMID: 30617339. Esteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115–118. 10.1038/nature21056 . Epub 2017 Jan 25. Erratum in: Nature. 2017;546(7660):686. doi: 10.1038/nature22985. PMID: 28117445; PMCID: PMC8382232. Rajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1:18. 10.1038/s41746-018-0029-1 . Aggarwal R, Sounderajah V, Martin G, Ting DSW, Karthikesalingam A, King D, Ashrafian H, Darzi A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. NPJ Digit Med. 2021;4(1):65. 10.1038/s41746-021-00438-z . PMID: 33828217; PMCID: PMC8027892. Yu AC, Mohajer B, Eng J. External Validation of Deep Learning Algorithms for Radiologic Diagnosis: A Systematic Review. Radiol Artif Intell. 2022;4(3):e210064. 10.1148/ryai.210064 . PMID: 35652114; PMCID: PMC9152694. Oakden-Rayner L, Dunnmon J, Carneiro G, Ré C. Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging. Proc ACM Conf Health Inference Learn (2020). 2020;2020:151–159. doi: 10.1145/3368555.3384468. PMID: 33196064; PMCID: PMC7665161. Hasanzadeh F, Josephson CB, Waters G, Adedinsewo D, Azizi Z, White JA. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit Med. 2025;8(1):154. 10.1038/s41746-025-01503-7 . PMID: 40069303; PMCID: PMC11897215. Food US, Administration D. Proposed regulatory framework for modifications to artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD): discussion paper and request for feedback. FDA; 2019. Available from: https://www.fda.gov/media/122535/download Sounderajah V, Ashrafian H, Golub RM, Shetty S, De Fauw J, Hooft L, STARD-AI Steering Committee, et al. Developing a reporting guideline for artificial intelligence-centred diagnostic test accuracy studies: the STARD-AI protocol. BMJ Open. 2021;11(6):e047709. 10.1136/bmjopen-2020-047709 . PMID: 34183345; PMCID: PMC8240576. He M, Li Z, Liu C, Shi D, Tan Z. Deployment of Artificial Intelligence in Real-World Practice: Opportunity and Challenge. Asia Pac J Ophthalmol (Phila). 2020 Jul-Aug;9(4):299–307. 10.1097/APO.0000000000000301 . PMID: 32694344. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, Moher D, Peters MDJ, Horsley T, Weeks L et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med. 2018;169(7):467–473. Available from: https://www.acpjournals.org /doi/10.7326/M18-0850 Zheng T, Lin F, Li X, Chu T, Gao J, Zhang S, Li Z, Gu Y, Wang S, Zhao F, Ma H, Xie H, Xu C, Zhang H, Mao N. Deep learning-enabled fully automated pipeline system for segmentation and classification of single-mass breast lesions using contrast-enhanced mammography: a prospective, multicentre study. EClinicalMedicine. 2023;58:101913. 10.1016/j.eclinm.2023.101913 . PMID: 36969336; PMCID: PMC10034267. Ng AY, Oberije CJG, Ambrózay É, Szabó E, Serfőző O, Karpati E, Fox G, Glocker B, Morris EA, Forrai G, Kecskemethy PD. Prospective implementation of AI-assisted screen reading to improve early detection of breast cancer. Nat Med. 2023;29(12):3044–9. 10.1038/s41591-023-02625-9 . Epub 2023 Nov 16. PMID: 37973948; PMCID: PMC10719086. Dembrower K, Crippa A, Colón E, Eklund M, Strand F, ScreenTrustCAD. Trial Consortium. Artificial intelligence for breast cancer detection in screening mammography in Sweden: a prospective, population-based, paired-reader, non-inferiority study. Lancet Digit Health. 2023;5(10):e703-e711. doi: 10.1016/S2589-7500(23)00153-X. Epub 2023 Sep 8. Erratum in: Lancet Digit Health. 2023;5(10):e646. 10.1016/S2589-7500(23)00181-4 . PMID: 37690911. Eisemann N, Bunk S, Mukama T, Baltus H, Elsner SA, Gomille T, Hecht G, Heywang-Köbrunner S, Rathmann R, Siegmann-Luz K, Töllner T, Vomweg TW, Leibig C, Katalinic A. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat Med. 2025;31(3):917–24. 10.1038/s41591-024-03408-6 . Epub 2025 Jan 7. PMID: 39775040; PMCID: PMC11922743. Salim M, Liu Y, Sorkhei M, Ntoula D, Foukakis T, Fredriksson I, Wang Y, Eklund M, Azizpour H, Smith K, Strand F. AI-based selection of individuals for supplemental MRI in population-based breast cancer screening: the randomized ScreenTrustMRI trial. Nat Med. 2024;30(9):2623–30. 10.1038/s41591-024-03093-5 . Epub 2024 Jul 8. PMID: 38977914; PMCID: PMC11405258. Glissen Brown JR, Mansour NM, Wang P, Chuchuca MA, Minchenberg SB, Chandnani M, Liu L, Gross SA, Sengupta N, Berzin TM. Deep Learning Computer-aided Polyp Detection Reduces Adenoma Miss Rate: A United States Multi-center Randomized Tandem Colonoscopy Study (CADeT-CS Trial). Clin Gastroenterol Hepatol. 2022;20(7):1499–e15074. Epub 2021 Sep 14. PMID: 34530161. Lachter J, Schlachter SC, Plowman RS, Goldenberg R, Raz Y, Rabani N et al. Novel artificial intelligence–enabled deep learning system to enhance adenoma detection: a prospective randomized controlled study. iGIE. 2023;2(1):52–8. Available from: https://www.igiejournal.org/article/S2949-7086(23)00015-8/fulltext Park DK, Kim EJ, Im JP, Lim H, Lim YJ, Byeon JS, Kim KO, Chung JW, Kim YJ. A prospective multicenter randomized controlled trial on artificial intelligence-assisted colonoscopy for enhanced polyp detection. Sci Rep. 2024;14(1):25453. 10.1038/s41598-024-77079-1 . PMID: 39455850; PMCID: PMC11512038. Lagström RMB, Bräuner KB, Bielik J, Rosen AW, Crone JG, Gögenur I, Bulut M. Improvement in adenoma detection rate by artificial intelligence-assisted colonoscopy: Multicenter quasi-randomized controlled trial. Endosc Int Open. 2025;13:a25215169. 10.1055/a-2521-5169 . PMID: 40018072; PMCID: PMC11866038. Dong L, He W, Zhang R, Ge Z, Wang YX, Zhou J, et al. Artificial Intelligence for Screening of Multiple Retinal and Optic Nerve Diseases. JAMA Netw Open. 2022;5(5):e229960. 10.1001/jamanetworkopen.2022.9960 . PMID: 35503220; PMCID: PMC9066285. Brant A, Singh P, Yin X, Yang L, Nayar J, Jeji D et al. Performance of a Deep Learning Diabetic Retinopathy Algorithm in India. JAMA Netw Open. 2025;8(3):e250984. doi: 10.1001/jamanetworkopen.2025.0984. Erratum in: JAMA Netw Open. 2025;8(4):e2511258. 10.1001/jamanetworkopen.2025.11258 . PMID: 40105843; PMCID: PMC11923701. Dai L, Sheng B, Chen T, Wu Q, Liu R, Cai C, et al. A deep learning system for predicting time to progression of diabetic retinopathy. Nat Med. 2024;30(2):584–94. 10.1038/s41591-023-02702-z . Epub 2024 Jan 4. PMID: 38177850; PMCID: PMC10878973. Marchetti MA, Cowen EA, Kurtansky NR, Weber J, Dauscher M, DeFazio J, Deng L, Dusza SW, Haliasos H, Halpern AC, Hosein S, Nazir ZH, Marghoob AA, Quigley EA, Salvador T, Rotemberg VM. Prospective validation of dermoscopy-based open-source artificial intelligence for melanoma diagnosis (PROVE-AI study). NPJ Digit Med. 2023;6(1):127. 10.1038/s41746-023-00872-1 . PMID: 37438476; PMCID: PMC10338483. Winkler JK, Blum A, Kommoss K, Enk A, Toberer F, Rosenberger A, Haenssle HA. Assessment of Diagnostic Performance of Dermatologists Cooperating With a Convolutional Neural Network in a Prospective Clinical Study: Human With Machine. JAMA Dermatol. 2023;159(6):621–7. 2023.0905. PMID: 37133847; PMCID: PMC10157508. Abdullah M, Sadia Afrin, Mridha MF, Sultan Alfarhood, Che D, Safran M. Explainable deep learning approaches for high precision early melanoma detection using dermoscopic images. Sci Rep. 2025;15(1). Trost B, Rodrigues L, Ong C, Dezellus A, Goldberg YH, Bouchat M et al. Artificial Intelligence Empowers Novice Users to Acquire Diagnostic-Quality Echocardiography. JACC: Advances. 2025;4(8):102005. Available from: https://www.sciencedirect.com/science/article/pii/S2772963X25004296 Lau ES, Di Achille P, Kopparapu K, Andrews CT, Singh P, Reeder C et al. Deep Learning-Enabled Assessment of Left Heart Structure and Function Predicts Cardiovascular Outcomes. Journal of the American College of Cardiology. 2023;82(20):1936–48. Available from: https://pubmed.ncbi.nlm.nih.gov/37940231/ Atwa AEM, Atlam ES, Ahmed A, Atwa MA, Abdelrahim EM, Siam AI. Diagnostics (Basel). 2025;15(15):1950. 10.3390/diagnostics15151950 . PMID: 40804914; PMCID: PMC12346745. Interpretable Deep Learning Models for Arrhythmia Classification Based on ECG Signals Using PTB-X Dataset. Bai X, Dong X, Li Y, Liu R, Zhang H. A hybrid deep learning network for automatic diagnosis of cardiac arrhythmia based on 12-lead ECG. Sci Rep. 2024;14(1):24441. 10.1038/s41598-024-75531-w . PMID: 39424921; PMCID: PMC11489693. Munjal P, Mahrooqi AA, Rajan R, Jeremijenko A, Ahmad I, Akhtar MI, Pimentel MAF, Khan S. Population-scale cross-sectional observational study for AI-powered TB screening on one million CXRs. NPJ Digit Med. 2025;8(1):418. 10.1038/s41746-025-01832-7 . PMID: 40634545; PMCID: PMC12241593. Sharma V, None Nillmani, Gupta S, Shukla KK. Deep learning models for tuberculosis detection and infected region visualization in chest X-ray images. Intell Med. 2023. Mujeeb Rahman KK, Zulaikha S, Dhafer B, Ahmed R. Advancing tuberculosis screening: A tailored CNN approach for accurate chest X-ray analysis and practical clinical integration. Intelligence-Based Med. 2025;11:100196. Ibrahim H, Liu X, Rivera SC, et al. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines. Trials. 2021;22:11. https://doi.org/10.1186/s13063-020-04951-6 . Tables Table 1 is available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files table1full24studies..xlsx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 14 Jan, 2026 Reviewers agreed at journal 08 Jan, 2026 Reviews received at journal 11 Dec, 2025 Reviewers agreed at journal 30 Nov, 2025 Reviewers agreed at journal 24 Nov, 2025 Reviewers agreed at journal 24 Nov, 2025 Reviewers invited by journal 05 Nov, 2025 Editor invited by journal 23 Oct, 2025 Editor assigned by journal 17 Oct, 2025 Submission checks completed at journal 17 Oct, 2025 First submitted to journal 16 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7876598","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Systematic Review","associatedPublications":[],"authors":[{"id":545448776,"identity":"b7016b2c-cc2c-41ca-b307-b1bf0d4e7256","order_by":0,"name":"Godswill Uzoechina","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABaElEQVRIie3Rv2vCQBQH8AuBy/Ks60n88S9cOIgK0o79N04CdglFKBQHoQGhXULn+F8EhM4HgbgIroodIgXp0MHiksFic1Uwim3XUvIlHJcjHy7vPYSyZPmboclTK+5eBCSLEpF4c4c15+tIcU4SAmmi0rqr8jyIHwjaE7lg0sEqL3j8JKlc9gYraCeEXM2jZfe5WNUEIhPA13T6Zr60UaPki1zwuifGKLzVQf4Y4ZbhhQuouxwZXg1u6MyuMg+1mC/OWrUU8WxTtbck1HNOAFRwZBEgij+zTR1Q0PQFmPSAsNWWNO/1D0nGEQrWmCr+dCTJ5phUiE31LbGwrkgy4YpDMG/2PZBESMKiVLsgqWWdEAwL1XBDSeY9VHAFy0NLlmmxfgBmumMPvcG7t26U81rSsbgbXNCxNYxJLMpYC5506JyXHocuW6ZuEbsN3p8dDEIWoSaTSt1yNKdvoi5//yZLlixZ/m0+AXMjegLcaSJXAAAAAElFTkSuQmCC","orcid":"","institution":"University of Nigeria Teaching Hospital (UNTH)","correspondingAuthor":true,"prefix":"","firstName":"Godswill","middleName":"","lastName":"Uzoechina","suffix":""},{"id":545448777,"identity":"f5c57909-198f-43f0-8869-75e44d7c7f3b","order_by":1,"name":"Treasure Osajiuba","email":"","orcid":"","institution":"University of Nigeria Teaching Hospital (UNTH)","correspondingAuthor":false,"prefix":"","firstName":"Treasure","middleName":"","lastName":"Osajiuba","suffix":""}],"badges":[],"createdAt":"2025-10-16 10:53:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7876598/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7876598/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96250681,"identity":"7fd12dba-34f3-40ee-8dd5-81ccae12d5c6","added_by":"auto","created_at":"2025-11-19 07:38:52","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":850549,"visible":true,"origin":"","legend":"","description":"","filename":"DeepLearninginClinicalDiagnosticsAScopingReviewofInnovationsShapingFutureHealthcareDelivery.docx","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/80dd9eea4d584d57c7b49300.docx"},{"id":96154996,"identity":"161fb416-d6f2-4abc-902a-ddeba376627f","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":52646,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.PRISMAFlowDiagram.docx","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/53916d07bc64e9304140089d.docx"},{"id":96155042,"identity":"407333af-d70f-4662-981f-7c9da25b6147","added_by":"auto","created_at":"2025-11-18 08:19:41","extension":"csv","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":24201,"visible":true,"origin":"","legend":"","description":"","filename":"table1full24studies.csv","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/4a15970e9f7c79b21b05a976.csv"},{"id":96251771,"identity":"b2ad0d9a-13aa-4714-925e-3676fb6cae69","added_by":"auto","created_at":"2025-11-19 07:40:00","extension":"json","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5580,"visible":true,"origin":"","legend":"","description":"","filename":"d4317a0a54b54529af10445be3d3b0c1.json","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/99ec38f81a292fe7af51d096.json"},{"id":96155000,"identity":"8447d632-ccff-40b9-a5e2-6f78ffb3ab66","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"xml","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":133211,"visible":true,"origin":"","legend":"","description":"","filename":"d4317a0a54b54529af10445be3d3b0c11enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/d91f30ff786799a204d1ce37.xml"},{"id":96155029,"identity":"b464ccf9-9324-4591-b218-db1f233e62cd","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"jpeg","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":787555,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/3438f5867a3782aedb4e1ee3.jpeg"},{"id":96155001,"identity":"60007403-931b-4379-a9ce-e78ef246c185","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":61332,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/9811722b1aed82df143b6b4d.png"},{"id":96155004,"identity":"0bff674f-f113-4f99-92bf-537e8059ee63","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":74622,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/a62b55556f5bcaa36c2d36ab.png"},{"id":96249256,"identity":"ec666c94-e752-498f-86c6-88b425b2ed62","added_by":"auto","created_at":"2025-11-19 07:31:11","extension":"jpeg","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":62534,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage12.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/982f4e175381339af1bff282.jpeg"},{"id":96155010,"identity":"9967f644-dcc8-4011-a304-41bf93a07c62","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpeg","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":628230,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage13.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/1dd5232eacbd9a2b05e8c3d8.jpeg"},{"id":96155006,"identity":"a85b82d3-8a4e-4ae1-970a-4eb3f26e10d4","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpeg","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36269,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage14.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/bf6f1caedce107cf873c1393.jpeg"},{"id":96155028,"identity":"a5df4166-629d-42d9-b217-6b5ca588f381","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"jpeg","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28721,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage15.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/3229b8127ee837c5d9df297a.jpeg"},{"id":96155008,"identity":"33a5748b-b58c-4c15-a4a6-7e8ac429b570","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpeg","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26301,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage16.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/798cff7ae4bf0c83e65d2453.jpeg"},{"id":96155007,"identity":"5c83e376-0a0d-405f-9b3b-7c42fa601cfb","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpeg","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1074,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage17.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/95ed10cb155b3bc8505561bf.jpeg"},{"id":96155014,"identity":"b9822f85-8276-481b-9944-de005bc88010","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpeg","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":962276,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/4bba9635ba558140418adb80.jpeg"},{"id":96248862,"identity":"e5414c84-27a3-4b3f-81d4-0d566192222b","added_by":"auto","created_at":"2025-11-19 07:29:31","extension":"jpeg","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":687696,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/e561fd6ca626c167156836e1.jpeg"},{"id":96155022,"identity":"b92d2b06-cffe-4c53-8a90-b1943786993c","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"jpeg","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":556431,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/4ff9d33b5dce02b4482045b4.jpeg"},{"id":96155036,"identity":"9a62dc0d-db46-4f7b-a1f5-f0b39b926008","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"jpeg","order_by":23,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":536895,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/71ccc93964502ae47195a903.jpeg"},{"id":96250671,"identity":"4aa2e756-03b0-4247-be1b-52c8dbd753ed","added_by":"auto","created_at":"2025-11-19 07:38:50","extension":"jpeg","order_by":24,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":582187,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/55ddf7a591bf0e3ab1088de7.jpeg"},{"id":96252023,"identity":"855ed3b9-0cf0-4819-aed3-e6769dedcff6","added_by":"auto","created_at":"2025-11-19 07:40:20","extension":"png","order_by":25,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36308,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/58f1717a14d5d0c54b267f97.png"},{"id":96250646,"identity":"01f18b42-3b09-4aa1-99a5-92d30cf71763","added_by":"auto","created_at":"2025-11-19 07:38:49","extension":"png","order_by":26,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":29896,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/14f2eb9c6b37b4d92b13836e.png"},{"id":96155026,"identity":"9d7f2f78-3e31-40aa-8813-d60de93e2874","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":27,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":93516,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/0073294642e2519412b683de.png"},{"id":96155012,"identity":"68492757-b023-49c0-8b3f-e152c85f2867","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":28,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":165140,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/76be26ea44b33e86729dcee4.png"},{"id":96155035,"identity":"3753e7de-7025-4523-8b02-6fc7272d5fa4","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":29,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":20271,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/ff13ab2cd6290dbffe1c0352.png"},{"id":96250829,"identity":"61f76d2e-5ce8-4c0d-a9e4-7ef1c02a2cff","added_by":"auto","created_at":"2025-11-19 07:39:03","extension":"png","order_by":30,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16025,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/c153454d6d17f963b367dd0f.png"},{"id":96155020,"identity":"a197ea9f-433e-4cd4-a95f-c67e1590991a","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":31,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11163,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage12.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/1df246ace70082b1ef896c42.png"},{"id":96250827,"identity":"4658d76c-fc5e-4ead-be05-17c049eb54b2","added_by":"auto","created_at":"2025-11-19 07:39:03","extension":"png","order_by":32,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":131089,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage13.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/adc76bce226f64e2dddaf65d.png"},{"id":96249777,"identity":"cba4a45c-8fd3-4b4b-8c52-92a70a46e6b4","added_by":"auto","created_at":"2025-11-19 07:36:13","extension":"png","order_by":33,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6586,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage14.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/634c16ee74a0e538ac9a4712.png"},{"id":96155037,"identity":"60a9737c-6b55-4bc4-84a1-7c03bfc991de","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":34,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5884,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage15.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/47a02aa33e0ff2573f4fc3d4.png"},{"id":96155024,"identity":"0a52ca0e-1791-44d3-a617-4b4a72c2df3e","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":35,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5061,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage16.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/b0dab52265d0c155f0ea58bf.png"},{"id":96155017,"identity":"405c0758-e87e-43be-82f4-e729ad5fdfd6","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":36,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":935,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage17.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/86b21aba3ad000dfe32d9d32.png"},{"id":96155033,"identity":"3c2277a4-c362-4c31-8cfb-b78a1bb4e9ce","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":37,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":204179,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/dbe48aaa01e15f3504c1b941.png"},{"id":96155015,"identity":"73a33bf4-2a5c-42d9-8926-36a8ac3af34c","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":38,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":141339,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/49d219042a27ae8c326350c1.png"},{"id":96251571,"identity":"09b991c9-10c0-4313-a400-379146e80ce2","added_by":"auto","created_at":"2025-11-19 07:39:49","extension":"png","order_by":39,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":118310,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/235a57c726cd5eb1c6a8dace.png"},{"id":96155019,"identity":"c3637c42-a132-45ff-a441-98b0d3f96c4e","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"png","order_by":40,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":113314,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/077c59c12129893b516a237a.png"},{"id":96249891,"identity":"21ed7dbd-3a14-4757-bf06-ac7196ac6cae","added_by":"auto","created_at":"2025-11-19 07:36:40","extension":"png","order_by":41,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":124069,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/79f85ef56caed390cc7d545b.png"},{"id":96250785,"identity":"904f1a0a-a911-4e3e-9002-188cbb428e9e","added_by":"auto","created_at":"2025-11-19 07:38:59","extension":"png","order_by":42,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12973,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/8289eb5ae089fa7c0d9c5eeb.png"},{"id":96155038,"identity":"0d010c1e-81ee-4151-9650-74fd8c29172e","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":43,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7141,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/49dad238ed234975df50620a.png"},{"id":96155027,"identity":"4f493168-cd04-4fd4-bcd4-0dd7e399bd87","added_by":"auto","created_at":"2025-11-18 08:19:21","extension":"png","order_by":44,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22909,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/5d836ef5c5dbdf8981095a73.png"},{"id":96248893,"identity":"38f9e883-c0b6-4742-a347-11bcb36fa0eb","added_by":"auto","created_at":"2025-11-19 07:29:34","extension":"xml","order_by":45,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":131600,"visible":true,"origin":"","legend":"","description":"","filename":"d4317a0a54b54529af10445be3d3b0c11structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/d3469caea381d20ce45ed359.xml"},{"id":96250638,"identity":"931e61df-fd0e-44b0-a115-8290453b6e60","added_by":"auto","created_at":"2025-11-19 07:38:49","extension":"html","order_by":46,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":145425,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/7b4e91487fb82e52cd2db416.html"},{"id":96154995,"identity":"7f007a8b-dd3d-46e3-864d-233c14eed892","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":117440,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePRISMA-ScR flowchart\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/6dc22bb26d33ddeb7c8d90a9.jpg"},{"id":96250925,"identity":"cb8dbba1-719f-46be-8b4b-2279790668c6","added_by":"auto","created_at":"2025-11-19 07:39:08","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":25606,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTimeline of publications by domain\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/1325ab674c3c997d09a8982d.jpg"},{"id":96155043,"identity":"0e01a9bb-1a22-45f9-8812-761b39d4e587","added_by":"auto","created_at":"2025-11-18 08:19:49","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":19494,"visible":true,"origin":"","legend":"\u003cp\u003eIllustration of the proportion of studies achieving external validation, prospective testing, RCT evidence, or post-deployment evaluation\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/b1a1b37aa830eb603e08634a.jpg"},{"id":96154997,"identity":"fcccaffb-d5a3-4d8a-9999-a592c8eaa746","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":29901,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNumber of Included Studies by Country\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/59056a7ae05a547207db2afc.jpg"},{"id":96155041,"identity":"26fa6bc2-bc6d-4f1c-9c98-7adfc07bc522","added_by":"auto","created_at":"2025-11-18 08:19:36","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":32842,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution of Deep Learning Model Families Across Clinical Domains with Representative AUC Values\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/760e88f87e5781fa5d2e0ba7.jpg"},{"id":96155040,"identity":"4bf3d98c-5cb8-41fe-8d6a-231488e0f6e3","added_by":"auto","created_at":"2025-11-18 08:19:26","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":27736,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFlow of Dataset Sources Through Model Families to Validation Stages\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/c18f60581dd3d287b78c5b10.jpg"},{"id":96257138,"identity":"ecab99b2-7ac2-41e8-9eed-aa981e39d340","added_by":"auto","created_at":"2025-11-19 07:51:36","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1357330,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/57c8baa1-3ed7-4fa7-8b8c-816f7726812d.pdf"},{"id":96155002,"identity":"f9a8ea25-bd39-4e35-beff-32f38b67b316","added_by":"auto","created_at":"2025-11-18 08:19:20","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":17857,"visible":true,"origin":"","legend":"","description":"","filename":"table1full24studies..xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7876598/v1/76900c9c210233bb8ef9b3fe.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deep Learning in Clinical Diagnostics: A Scoping Review of Innovations Shaping Future Healthcare Delivery","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDeep learning (DL), a subset of machine learning based on representation learning through multilayer neural networks, has rapidly gained traction in clinical diagnostics, with the potential for automated image interpretation, real-time procedure guidance, signal classification, and risk-stratified screening on a large scale (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Landmark demonstrations in dermatology and other image-rich domains have provided evidence that DL can achieve expert-level performance on specialized tasks, while large-scale endeavors focused on applying DL to electronic health records and population\u0026ensp;screening demonstrate applicability across modalities and data types (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAlthough substantial accuracies are reported in many single-study reports, the translational pathway from algorithmic development to safe, effective clinical implementation has been uneven. Systematic reviews describe significant heterogeneity in study designs, lack of consistent external validation, and pervasive differences in reporting standards; factors that can inflate apparent performance and limit confidence in real-world applicability (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e5\u003c/span\u003e). The divide between retrospective development and operational deployment is further widened by datasets limitations that allow for hidden stratification and other modes of failure: models that have strong average performance may still fail systematically on clinically relevant but under-represented subgroups or imaging subsets (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eBias, fairness, and robustness concerns have led to renewed focus on dataset auditing, subgroup\u0026ensp;reporting, and other formal procedures for identifying and addressing distributional and label-quality problems (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e7\u003c/span\u003e). Regulators and governing bodies have also hinted towards the requirement for lifecycle methodologies for AI/ML medical software, emphasising pre-specified change-control plans, ongoing performance monitoring, and transparent evidence of clinical safety as necessary conditions for deployment (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e8\u003c/span\u003e). At the same time, reporting guidelines for diagnostic AI are being developed to define a standard way to report accuracy, thresholds, calibration, and implementation settings to allow clinicians, payors, and regulators to better assess readiness for use (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eCollectively, these changes thus present an urgent demand for mapping the current evidence landscape: which clinical fields have advanced beyond retrospective development to prospective evaluation or real-world implementation, what range of potential outcomes beyond diagnostic accuracy (such as workflow impact, safety indicators, or clinician-assistance advantages) have been assessed, and which cross-cutting methodological and implementation challenges recur across studies (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e10\u003c/span\u003e). The scoping review aims to systematically map published DL diagnostic studies, classify their validation and implementation phase, and integrate operational and equity findings to guide researchers, clinicians, and policy makers on where DL is ready for clinical support and where further work remains.\u003c/p\u003e"},{"header":"Methodology","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eStudy Design\u003c/h2\u003e\u003cp\u003eWe performed a scoping review to map the current literature on deep-learning (DL) based diagnostic systems and assess their clinical validation\u0026ensp;and implementation. The review was conducted in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews) statement to ensure transparent and reproducible reporting of objectives, search methodology, selection, and data extraction (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e11\u003c/span\u003e). The review protocol, which outlined the inclusion\u0026ensp;criteria, search strategy, and data-charting elements, was established in advance.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eEligibility criteria\u003c/h3\u003e\n\u003cp\u003eWe considered primary empirical studies published in peer-reviewed journals, which described the development, validation, or clinical application of a deep learning based model for a diagnostic task in human healthcare. Eligible study designs included retrospective development study, prospective evaluation,\u0026ensp;randomized controlled trial, quasi-randomized trial, post-deployment audit, or real-world application report. We considered studies across modalities (imaging, signals, endoscopy, dermoscopy, ECG,\u0026ensp;etc.) so long as the studies reported at least one quantitative diagnostic or implementation outcome (e.g., AUC, sensitivity, specificity, detection rate, ADR, workflow/time metrics). Excluded were\u0026ensp;review articles, editorials, purely technical papers with no clinical data, animal research, and studies applying only classical (non-deep) machine learning methods. We restricted ourselves to English-language full texts. The eligible time window was 2022\u0026ndash;2025 to cover recent\u0026ensp;DL methodologies that could be translated into clinical applications.\u003c/p\u003e\n\u003ch3\u003eInformation sources and search strategy\u003c/h3\u003e\n\u003cp\u003eWe conducted a search of PubMed and Google Scholar for articles published between 1 January 2020 and 30 September 2025. The search combined MeSH terms and free text terms relating to three concepts: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) deep learning/convolutional neural networks/transformer; (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2\u003c/span\u003e) diagnosis/screening/detection/classification; and (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e3\u003c/span\u003e) validation/implementation/trial/deployment. An illustration\u0026ensp;search string for PubMed was:\u003c/p\u003e\u003cp\u003e(\"deep learning\" OR \"convolutional neural network\" OR \"CNN\" OR \"transformer\" OR \"neural network\") AND (\"diagnosis\" OR \"screening\" OR \"detection\" OR \"classification\") AND (\"validation\" OR \"prospective\" OR \"randomized\" OR \"trial\" OR \"deployment\" OR \"implementation\").\u003c/p\u003e\u003cp\u003eSecondary search methods included manual handsearching, including backward and forward citation tracking of eligible studies, and screening the reference lists of relevant reviews\u003c/p\u003e\n\u003ch3\u003eStudy selection\u003c/h3\u003e\n\u003cp\u003eAll results obtained\u0026ensp;were loaded into a citation manager, and duplicates were removed. Title/abstract screening was conducted by two independent reviewers using pre-defined inclusion criteria; potentially relevant full texts were subsequently retrieved for a duplicate assessment. Differences in opinion at any screening level were resolved by consensus; further disagreements were mediated by a third reviewer. We recorded the full-text exclusions with reasons for exclusion, and the selection flow is presented as a\u0026ensp;PRISMA-ScR flow diagram (Fig.\u0026nbsp;1).\u003c/p\u003e\n\u003ch3\u003eData charting\u003c/h3\u003e\n\u003cp\u003eWe created a structured data spreadsheet in Microsoft Excel for data extraction and charting, which corresponded to the columns of the provided extraction table. From each study, we extracted the following details verbatim where possible:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eAuthor(s) and year of publication; DOI/URL if available;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eCountry; Clinical domain; Diagnostic task/modality;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSample size (N) and dataset source;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eDeep learning model type (architecture/ensemble/commercial vendor name);\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eComparator (human readers, standard of care, other algorithms);\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eOutcome\u0026ensp;metrics reported (AUC, sensitivity/specificity, detection rates, DSC, MAE, reading time, workload metrics, etc.).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eValidation approach (internal\u0026ensp;train/validation/test, external test, temporal split, prospective test, randomized design, post-deployment audit);\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSetting of Application (single centre, multicenter, screening program, deployment center);\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMain Findings (primary interpretations of authors);\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eReported Challenges/Limitations.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eOne reviewer conducted the initial extraction, and a second reviewer independently verified all extracted fields for completeness and accuracy. Discrepancies were resolved by discussion and, when necessary, review of the full text. Figure\u0026nbsp;1. PRISMA-ScR flow diagram summarising study selection. From an initial 2,134 records, 24 studies met the inclusion criteria and were included in the final scoping review.\u003c/p\u003e\u003cp\u003e\u003cb\u003eFigure 1: PRISMA-ScR flowchart\u003c/b\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cem\u003eStudy selection\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eWe identified the body of literature that met our pre-specified inclusion criteria through both database and grey-literature searches (see Methods). Following the removal of duplicates and screening of titles/abstracts, full texts were assessed for eligibility, and findings were mapped for studies that reported a deep-learning method applied to a human clinical diagnostic task. The study selection process is summarized in the PRISMA-ScR flow diagram (see Methodology)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eStudy characteristics\u0026nbsp;\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe included studies covered various clinical domains, with study designs ranging from retrospective development studies to large prospective deployments and randomized trials. Table 1 reports a summary of the study domains, sample size, dataset sources, deep learning architectures, validation methods, and the stage of clinical translation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1: Study characteristics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eClinical domains most frequently highlighted were radiology (particularly breast imaging and chest radiography), gastroenterology (colonoscopy polyp detection), ophthalmology (fundus-based screening and prognostics), dermatology (dermoscopy), and cardiology (echocardiography and ECG). Pulmonology (TB detection on CXR) was also emerging as a field with large datasets and post-deployment evaluations.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe size of the datasets ranged from small clinical studies (hundreds of patients/images) to very large screening datasets (hundreds of thousands to over one million images in deployment settings). Model types covered traditional CNN variants (ResNet, Xception, EfficientNet), ensemble techniques, multi-task models with temporal transformers (echocardiography), object-detection architectures (RetinaNet), as well as explainable/gradient-based visualizations (Grad-CAM). A few of the studies reported commercially available systems (e.g., Lunit, ScreenTrustCAD, GI Genius, ARDA), while others used open-source or custom models.\u003c/p\u003e\n\u003cp\u003eValidation methodologies varied from retrospective training/testing splits to multi-center external validation, prospective trials, randomized controlled trials, and extensive post-deployment audit studies.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable 2 summarizes studies by clinical domains reporting counts, common diagnostic tasks, median/mean sample size, and median primary performance metric (e.g., AUC or ADR where applicable) for each clinical domain.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2: Summary by clinical domain.\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eClinical domain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNumber of included studies (n)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCommon diagnostic tasks (typical modalities)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian sample size (per study) \u0026mdash; range (min\u0026ndash;max)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCommon primary performance metric(s) reported\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian performance (metric + value) \u0026mdash; basis (n studies used)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNotes/interpretation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRadiology: Breast imaging\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eScreening mammography (AI as additional/independent reader); CEM lesion segmentation/classification; AI triage to MRI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e40,062 (range 1,912 \u0026ndash; 463,094)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCancer detection rate (CDR) / relative % change; AUC for lesion classification (when reported)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian relative increase in cancer detection \u0026asymp; 13.8% (computed across studies that reported relative or % change: Eisemann +17.6%, Chang +13.8%, Dembrower +4% \u0026rarr; median 13.8%, n=3). Single-study AUC example: Zheng et al. AUC \u0026asymp;0.94 (internal/external).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStrong translation activity: several large prospective/rollout studies and RCTs; heterogeneous metrics (detection rates, PPV, recall) so use domain-specific metrics rather than pooled AUC.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGastroenterology: Colonoscopy (CADe)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eReal-time polyp/adenoma detection (video/white-light)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e735 (range 223 \u0026ndash; 805)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAdenoma detection rate (ADR), Polyp detection rate (PDR), Adenomas per colonoscopy (APC), miss rates\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian ADR with AI \u0026asymp; 43.7% (AI ADRs: 50.44%, 37%, 35%, 59.1% \u0026rarr; median 43.72%, n=4). Median control ADR \u0026asymp; 35.8% (controls median 35.82%, n=4). Median absolute ADR improvement \u0026asymp; 7.9 percentage points (n=4).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eConsistent RCT evidence of clinically meaningful ADR/PDR improvements across multiple centers; median ADR improvement ~8 pp. FP alerts and withdrawal time reported variably.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDermatology: Dermoscopy / Melanoma detection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDermoscopy-based melanoma vs benign lesion classification; clinician + AI assistance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e435 patients (range 228 lesions \u0026ndash; 13,900 images)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAUC; sensitivity / specificity; accuracy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian AUC \u0026asymp; 0.945 (AUC values used: 0.816; 0.9211; 0.968; 0.9868 \u0026mdash; median \u0026asymp; 0.945, n\u0026asymp;4 AUC values from 3 studies). Typical reported sensitivity (with AI) often high (\u0026asymp;95\u0026ndash;97%) (n studies vary).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHigh sensitivity in many studies but specificity variable (trade-offs). Dataset diversity (skin tones) limited in several studies and many results are dataset/computational rather than large prospective deployments.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eOphthalmology: Fundus screening \u0026amp; prognostics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMulti-disease fundus screening; diabetic retinopathy screening; prognostic (time-to-progression) models\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e18,760 (range 4,537 \u0026ndash; 110,784) \u0026mdash; per-study sizes (individuals/images)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eScreening: sensitivity/specificity; Prognostics: C-index (survival)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian screening sensitivity \u0026asymp; 91.4% (screening studies: Dong 89.8%, Brant 97.0%, Paisan 91.4% \u0026rarr; median 91.4%, n=3). Prognostic (DeepDR Plus) median C-index \u0026asymp; 0.84 (external range 0.823\u0026ndash;0.862; use midpoint \u0026asymp; 0.84, n=1 study with multiple cohorts).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMix of massive population screening deployments (high external validity) and prognostic models enabling personalized screening intervals. Strong external validations but geographic training concentration (China/India) noted as a limitation.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePulmonology / Chest radiography: TB \u0026amp; general CXR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eTB detection on chest X-ray; segmentation/localization (Grad-CAM)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e12,848 (range ~2,000 \u0026ndash; 1,040,000)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAUC; accuracy; false-negative rate (FNR) / workload reduction (WLR)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian AUC \u0026asymp; 0.992 (AUCs: Sharma 0.999; Munjal 0.9851 \u0026rarr; median \u0026asymp; 0.992, n=2 AUCs). Other studies report extremely high accuracy/accuracy ~99\u0026ndash;100% (Mujeeb).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSome studies are very large real-world deployments (Munjal ~1.04M CXRs) demonstrating operational safety and large workload reductions; many high AUC/accuracy numbers come from public datasets or retrospective analyses \u0026mdash; prospective generalizability needs ongoing evaluation.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCardiology: Echocardiography (TTE)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAutomated full TTE interpretation (view classification, quantification)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e48,146 (median of reported per-study values: 32,265 and 64,028 \u0026rarr; median 48,146; range 32,265 \u0026ndash; 64,028)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAUC for classification tasks; MAE for quantitative estimates (LVEF MAE)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian classification AUC \u0026asymp; 0.94 (Holste median AUC 0.91; Emily S. view AUROC \u0026gt;0.97 \u0026rarr; median \u0026asymp; 0.94, n=2 studies). LVEF MAE \u0026asymp; 4.2\u0026ndash;4.5% reported in examples.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eVery large video datasets and multi-task models with external validations; results show near-expert performance for many labels and good quantification accuracy, but prospective outcome studies and workflow integration remain limited.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCardiology: ECG arrhythmia classification\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e12-lead ECG automated arrhythmia detection / multi-class classification\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e10,943 (median of 48 (MIT-BIH) and 21,837 (PTB-XL) \u0026rarr; 10,943; range 48 \u0026ndash; 21,837)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAccuracy; AUC (when reported)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMedian accuracy / AUC \u0026asymp; high (\u0026asymp; 98%) \u0026mdash; examples: Bai (MIT-BIH / PTB generalization) accuracy/AUC ~99.4% / 0.9875; Atwa (PTB-XL) multiclass accuracy typically 71\u0026ndash;98% depending on task \u0026mdash; median across available figures \u0026asymp; ~98% (binary/superclass) (n=2 studies; metric differs by classification granularity).\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBenchmark datasets show near-perfect classification on some tasks, but MIT-BIH is small and older \u0026mdash; real-world generalizability requires larger clinical dataset validation and prospective testing.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFigure 2 highlights annual counts of included studies, stacked by clinical domain, demonstrating the temporal growth and domain shifts in DL diagnostic research.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 2: Timeline of publications by domain\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eThematic synthesis by clinical domain\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRadiology: Breast imaging and mammography screening\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA significant portion of the included studies assessed DL for breast imaging, with works on automated lesion segmentation, single-read support, AI as independent readers in population screening, and AI triage for supplemental MRI. Representative findings include:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eSegmentation \u0026amp; single-mass classification (contrast-enhanced mammography): Zheng et al., 2023 proposed and evaluated a Fully Automated Pipeline System (RefineNet + Xception + Pyramid Pooling) for a cohort of 1,912 women and achieved Dice similarity coefficients of 0.888 (internal), 0.820 (external), prospective 0.837, and AUC of 0.947 (internal) and 0.940 (external); the system decreased reading time (6 seconds per case vs ~3 minutes) and enhanced radiologist performance and BI-RADS reclassification in ~12\u0026ndash;13% of\u0026ensp;the cases. The study demonstrates\u0026ensp;significant segmentation/classification accuracy as well as practical workflow benefits (12).\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eAI as an additional or replacement reader in population screening: Multiple large prospective and real-world studies highlight that AI can safely augment or replace one human reader:\u0026nbsp;\u003col\u003e\n \u003cli\u003eNg A.Y. et al., 2023 (Hungary) stated that the integration of an ensemble commercial AI (\u0026quot;Mia\u0026quot;) led to an increase in cancer detection by 0.7\u0026ndash;1.6 per 1,000 and a slight recall increase (0.16\u0026ndash;0.30%). The study matured from pilot to live rollout and demonstrated feasibility in a screening program (13).\u003c/li\u003e\n \u003cli\u003eDembrower et al., 2025 (Sweden) screened 55,581 women and demonstrated that AI alone or displacing one radiologist with AI was non-inferior in cancer detection rate compared to two-radiologist reads; triple reading (2 radiologists + AI) increased cancer detection by 8% (14).\u003c/li\u003e\n \u003cli\u003eEisemann et al, 2025\u0026ensp;(Germany), AI-assisted double reading significantly increased cancer detection (6.7\u0026ndash;5.7 per 1,000; +17.6%), and biopsy PPV also improved, across 463,094 screens performed (15).\u0026nbsp;\u003c/li\u003e\n \u003c/ol\u003e\n \u003c/li\u003e\n \u003cli\u003eAI triage for MRI: Salim et al., 2024 applied AI to triage the top ~6.9% for adjunct MRI, detecting 64.4\u0026ensp;cancers/1,000 MRIs as opposed to 16.5/1,000 with density-based triage; a nearly 4\u0026times; increase in yield\u0026ensp;per MRI (16).\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eMany of the breast imaging studies are prospective, population-based, or linked to national programs and\u0026ensp;are among the highest levels of translation in our corpus (see Table 3 and Figure 4). A common limitation is\u0026ensp;the restricted geographic variability (many are from Sweden, South Korea, Hungary, and China) and reliance on a few commercial systems. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable 3 reports counts and proportions of studies in each translation stage: development only, internal validation, external validation, prospective clinical testing, randomized trials, and post-deployment surveillance/real-world audits.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3: Deployment readiness across included studies\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTranslation stage (highest achieved)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eCount (n)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e% of studies\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eExample studies (representative; author, year)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eProspective clinical testing (prospective evaluations, paired-reader, prospective cohort studies, prospective external validation)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e8\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e33.3%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eZheng T et al., 2023; Dong L et al., 2022; Dembrower K et al., 2025; Chang Y-W et al., 2025; Marchetti MA et al., 2023; Winkler JK et al., 2023; Yifan Peng et al., 2024; Paisan Ruamviboonsuk et al., 2022\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRandomized trials (RCTs and quasi-randomized trials)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e20.8%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGlissen Brown et al., 2022 (RCT); Park D.K. et al., 2024 (RCT); Latcher et al., 2023 (RCT); Salim M. et al., 2024 (RCT); Lagstr\u0026ouml;m RMB et al., 2025 (quasi-randomized)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePost-deployment surveillance / real-world audits / large rollouts\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e16.7%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNg A.Y. et al., 2023 (rollout); Munjal P. et al., 2025 (large screening deployment); Brant A. et al., 2025 (post-deployment adjudication); Eisemann N. et al., 2025 (multi-site real-world implementation)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eExternal validation (multi-center or multi-cohort external testing, but not prospective clinical testing)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e8.3%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHolste G. et al., 2025; Emily S., 2023\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDevelopment only / computational (retrospective train/val/test, public-dataset benchmarks)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e20.8%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMahmud M.A.A., 2025; Sharma V. et al., 2024; Mujeeb et al., 2024; Xiangyun Bai, 2024; Ahmed E.M. Atwa et al., 2025\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e24\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e100%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eGastroenterology: Colonoscopy polyp/adenoma detection (CADe)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMultiple randomized and quasi-randomized trials reported clinically meaningful rises in adenoma detection rate (ADR), reductions in miss rates, and increases in polyps detected per colonoscopy (APC). Representative findings are:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eRandomized tandem or parallel RCTs:\u003col\u003e\n \u003cli\u003eGlissen Brown et al., 2022 (US, EndoScreener, SegNet) in 223 patients decreased adenoma miss rate (20.12% vs 31.25%, P=0.0247) as well as polyp miss rate (20.7% vs 33.7%). First-pass ADR was trended higher with CADe (17).\u003c/li\u003e\n \u003cli\u003eLatcher et al., 2023 (Israel, DEEP\u0026sup2; CADe) reported ADR 37% vs 27% (+10%, p=0.0057) and PDR 55.5% vs 38.7%\u0026ensp;(p\u0026lt;0.001) (18).\u003c/li\u003e\n \u003cli\u003ePark D.K. et al., 2024 (Korea, RetinaNet) revealed PDR 62% vs 52% (p=0.01) and ADR 35% vs 28% (p=0.03) (19).\u0026nbsp;\u003c/li\u003e\n \u003c/ol\u003e\n \u003c/li\u003e\n \u003cli\u003eLarge quasi-randomized trial / real-world rollout:\u003col\u003e\n \u003cli\u003eLagstr\u0026ouml;m RMB et al., 2025 (Denmark, GI Genius v3) in 795 patients highlighted ADR 59.1% vs 46.6% (p\u0026lt;0.001) (20).\u003c/li\u003e\n \u003c/ol\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eCADe systems usually operate in real-time, considering false-positive (FP) alert rates and withdrawal time. Many studies report a modest rise in procedure time or no change;\u0026ensp;FP alerts per exam are sometimes quantified (e.g., ~4 false alerts/exam - Latcher et al.). The body of randomized evidence ranks colonoscopic CADe among the most solidly validated DL diagnostic applications.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOphthalmology: Fundus-image screening and prognostics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOphthalmology studies included multi-disease classifiers for screening and DL prognostic models that predict time-to-progression. Finings are:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eMulti-disease screening (RAIDS): Dong et al., 2022 trained RAIDS on 120,002 development images and prospectively validated its performance on 208,758 images (110,784 individuals). RAIDS sensitivity for any of the 10 retinal diseases was 89.8% with per-disease accuracy ranging up to 95\u0026ndash;99.9%; RAIDS was comparable or better than human graders for several diseases in multicenter screening practice (21).\u0026nbsp;\u003cbr\u003e\u0026nbsp;\u0026nbsp;\u003c/li\u003e\n \u003cli\u003ePost-deployment audit (ARDA): Brant et al., 2025 assessed ~4,537 adjudicated images sampled from ~600,000 screened images from the Aravind program and demonstrated\u0026ensp;severe NPDR/PDR sensitivity at 97.0% and specificity at 96.4%, indicating robust post-deployment performance for a CE-approved ensemble (22).\u003c/li\u003e\n \u003cli\u003ePrognostics: DeepDR Plus (Yifan Peng et al., 2024), a model trained on 717,308\u0026ensp;images, externally validated on 8 longitudinal cohorts, reporting C-indices for 1\u0026ndash;5 year DR progression prediction of ~0.82\u0026ndash;0.86; a prospective real-world validation demonstrated safely prolonged screening intervals with minimal delayed detection (23).\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eOphthalmology includes both large prospective screening deployments and prognostic models that have been\u0026ensp;externally validated. Authors\u0026rsquo; concerns relate to generalizability to areas beyond the regions where they were trained (mainly China/India) and reliance on image quality.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDermatology: Dermoscopy and melanoma detection\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA few prospective or diagnostic results on CNN assistance for dermoscopic melanoma classification were obtained in\u0026ensp;dermatology: they revealed high sensitivity, but only moderate specificity. Findings were:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eMarchetti et al., 2023\u0026ensp;(ADAE): prospective observational study (603 biopsied lesions, 95 melanomas) in which ADAE obtained a sensitivity of 96.8% and increased dermatologist AUC from 0.780 \u0026rarr; 0.816 (p=0.042); however, specificity was still low (37.4%) (24).\u003cbr\u003e\u0026nbsp;\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eWinkler et al., 2023\u0026ensp;(Moleanalyzer Pro): sensitivity of the dermatologist was increased from 84.2% to 100% by CNN assistance in 228 lesions; specificity and accuracy also improved (25).\u003c/li\u003e\n \u003cli\u003eMahmud et al., 2025 reported high accuracy and AUC across large public dermoscopy datasets using an Xception-based model with explainability, but this was computational (no clinical deployment) and raised concerns about skin-tone representation (26).\u003cbr\u003e\u0026nbsp;\u003cbr\u003e\u003cstrong\u003eCardiology: Echocardiography and ECG\u003c/strong\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eStudies in cardiology ranged from high-volume automated transthoracic echocardiogram (TTE) interpretation to ECG arrhythmia detection. Representative findings were:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eEchocardiography (PanEcho): Trost et al., 2025 trained a model on ~1.2 million echo videos and reported a median classification AUC\u0026ensp;of 0.91 (IQR 0.88\u0026ndash;0.93) and LVEF MAE \u0026asymp;4.2\u0026ndash;4.5% (internal/external), indicating multi-task automation across numerous labels and extension to abbreviated or POC echo (27).\u003c/li\u003e\n \u003cli\u003eAutomated quantification and prognostic models: Emily S., 2023 (DROID workflow) reported AUROCs \u0026gt;0.97 for\u0026ensp;view classification and showed prognostic associations (e.g., HR 1.43 per 1 SD decrease in model-derived LVEF for heart failure) (28).\u003c/li\u003e\n \u003cli\u003eECG arrhythmia classification: Atwa et al., 2025 and Xiangyun Bai, 2024 achieved very high accuracies in benchmark datasets (MIT-BIH, PTB, PTB-XL), with hybrid CNN + attention approaches reporting \u0026gt;99% in certain binary cases and promising multiclass results (29,30).\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003ePulmonology: Chest X-ray (TB screening and broader chest CXR)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePulmonology studies included large-scale screening applications as well as cross-dataset validations for TB identification. Findings were:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eLarge-scale deployment (AIRIS-TB): Munjal et al., 2025 evaluated AIRIS-TB on ~1.04 million CXRs with an AUC of 98.51%\u0026ensp;and TB-FNR of \u0026nbsp; 0% (safe setting), reporting up to 80% automated reporting capability without compromising safety (31).\u003c/li\u003e\n \u003cli\u003eOther high-volume screening: Sharma et al., 2024, and Mujeeb et al., 2024 reported very high accuracies for TB detection using UNet segmentation + Xception classification and lightweight CNNs, respectively; however, they were constrained to public datasets or smaller curated sets (32,33).\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cem\u003eCross-cutting themes and methodological observations\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eFirst, with respect to the rigor of validation and readiness for translation, several clinical domains, most notably breast cancer screening, colonoscopy computer-aided detection (CADe), and some ophthalmology screening programs, demonstrate mature levels of clinical validation. These include rigorously designed randomized controlled\u0026ensp;trials (RCTs) and even large-scale national rollouts. Notable examples include Dembrower et al. (2025), Eisemann et al. (2025), Lagstr\u0026ouml;m et al. (2025), and Ng et al. (2023), which demonstrated that deep learning (DL) systems can be\u0026ensp;successfully implemented in real-life clinical environments. But many studies remain confined to the retrospective computational phase, relying primarily on internal or public datasets with conventional train-validation-test splitting and lacking external validation or clinical evaluation. Typical examples are Mahmud (2025) in dermoscopy and Mujeeb et al. (2024) in tuberculosis chest X-ray (CXR) classification. This is visualised in Figure 3, which shows the proportion of studies that attained external validation, prospective testing, RCT evidence, or post-deployment evaluation, and also reveals differences in translation readiness across domains.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFigure 3: Illustration of the proportion of studies achieving external validation, prospective testing, RCT evidence, or post-deployment evaluation\u003cbr\u003e\u0026nbsp;Concerning explainability and interpretability practices, the majority of the studies featured qualitative explainability tools such as Grad-CAM, saliency maps, or confidence overlay, especially in dermoscopy, chest radiography, and TB imaging (Mahmud 2025; Sharma\u0026ensp;2024; AIRIS-TB). Nevertheless, these visualizations were often expressed in a purely descriptive manner without structured human-factors evaluation or quantitative assessment of interpretability. In studies where explainability techniques were tightly coupled with clinical workflows (e.g., Grad-CAM-based overlays utilised by radiologists or dermatologists), authors noted that such tools facilitated clinician confidence and transparency, without supplanting human adjudication or expert oversight.\u003c/p\u003e\n\u003cp\u003eA persistent deficiency was identified in terms of dataset diversity, fairness, and demographic reporting (Figure 4 highlights the number of included studies by country). Many studies neglected to provide detailed reports on the demographic composition of the datasets, including race, ethnicity, skin tone, or socio-economic status, or on the performance metrics of specific subgroups. This limitation was significantly evident in dermatology and some radiology studies. In the relatively few cases where subgroup analyses were conducted, performance often changed substantially between demographic subgroups, highlighting the limited generalizability of the model and the need for diverse, representative external validation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 4: Number of Included Studies by Country\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConcerning reporting quality and reproducibility, only a handful of studies rendered their source code or model weights publicly available. A lot of them relied on proprietary algorithms or commercial platforms, which hindered independent validation and objective benchmarking. Reporting practices were similarly inconsistent: while AUC (Area Under\u0026ensp;the Curve) was almost universally reported, some details, such as decision thresholds, calibration methods, and measures of uncertainty, were frequently unreported. Sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) at clinically relevant thresholds were rarely reported.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eImplementation outcomes \u0026amp; reported challenges\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eAmong the included studies, implementation-related outcomes mainly focused on the efficiency of operation, diagnostic performance, and safety, representing both the potential and challenges of real-world AI application in clinical practice. The operational outcomes reported indicated a considerable promise\u0026ensp;in accelerating workflows and redistributing workloads. For example, Zheng et\u0026ensp;al. (2023) estimated that the fully automated FAPS classification took about six seconds per case compared to nearly three minutes of radiologist review, illustrating significant time savings. In the same manner, AIRIS-TB deployment as reported by Munjal et al. (2025) achieved up to 80% automated reporting with sustained diagnostic safety. However, not all findings reflected workload reduction: Ng et al. (2023) noted a 4\u0026ndash;6% workload increase when implementing AI as a second reader in screening workflows, and Eisemann et al. (2025) found slight increases in reading volume along with enhanced cancer detection rates and positive predictive value.\u003c/p\u003e\n\u003cp\u003eFigure 5 illustrates the distribution of study counts across deep learning model families and clinical domains, showing both the frequency of utilization and corresponding mean diagnostic performance (AUC) within each category.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 5. Distribution of Deep Learning Model Families Across Clinical Domains with Representative AUC Values\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDiagnostic and management-related impacts were also described. Zheng et al. (2023) reported that AI-assisted mammographic interpretation led to BI-RADS category changes that altered management decisions in about 12\u0026ndash;13% of patients, highlighting AI\u0026rsquo;s real impact on downstream clinical care. In another case,\u0026ensp;Salim et al. (2024) reported a\u0026ensp;4-fold higher diagnostic yield per MRI scan using AI-based MRI triage. In addition to detection, prognostic tools, for instance DeepDR Plus (Yifan Peng et\u0026ensp;al., 2024), extended the scope of AI to outcome forecasting; predicting diabetic retinopathy progression risk and safely extended screening intervals with a \u0026le;0.18% risk rate of missing vision-threatening retinopathy at detection to a whole new level of outcome forecasting\u0026mdash;predicting diabetic retinopathy progression risk and safely extending screening intervals with a \u0026le;0.18% risk rate of delayed vision-threatening retinopathy detection. Taken together, these examples show that well-calibrated models can not only improve diagnostic efficiency but also optimize resource utilization and personalize screening strategies. Figure 6 presents a Sankey diagram highlighting how dataset sources link to different deep learning model families and their corresponding validation stages, illustrating the progression of data utilization from public or clinical origins to various levels of model evaluation and implementation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 6. Flow of Dataset Sources Through Model Families to Validation Stages\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDespite these encouraging results, several implementation challenges and barriers were invariably reported. Generalizability and external validity were flagged as significant concerns: many models had lower performance when tested in populations that differed geographically, ethnically, or in imaging modality. Many authors highlighted the importance of diverse, multi-ethnic external validation cohorts. Concerns about dataset and spectrum bias were similarly raised, stating that publicly available datasets or datasets curated by research groups may not truly capture the real-world heterogeneity of disease, leading to the inflation of performance estimates.\u003c/p\u003e\n\u003cp\u003eExplainability and clinician acceptance posed further barriers. While visualization methods (such as saliency or Grad-CAM maps) have improved the interpretability of the models, they are not sufficient on their own to establish full clinical trust, prompting suggestions for structured human-factors studies and workflow-integrated user interfaces. Infrastructural and operational difficulties were reported as well: how best to integrate PACS with EHR systems, computational latency for real-time analysis, and the challenge of instituting escalation pathways for AI-flagged cases were persistent issues in many of the reports.\u0026nbsp;\u003cbr\u003e\u0026nbsp;\u003cbr\u003e\u0026nbsp;Finally, the considerations for regulation and ethics were uneven. Many studies involved commercial algorithms or industry collaboration, but inconsistently reported conflicts of interest, and regulatory status such as CE marking, FDA clearance. Resource-related constraints were especially relevant in low- and middle-income environments, where the cost of hardware and connectivity played a role in shaping the decisions around model design. Several teams explicitly focused on lightweight or edge-compatible models, for example, Mujeeb et al. (2024), who optimized for constrained computational infrastructure, to enable equitable.\u003c/p\u003e"},{"header":"Discussions","content":"\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003ePrincipal Findings\u003c/h2\u003e\u003cp\u003eThis scoping review demonstrated that current deep-learning (DL) diagnostic research varies widely in both maturity and the quality of evidence. Certain domains, such as screening mammography, real-time colonoscopy CADe, and large-scale chest X-ray tuberculosis (TB) screening, have moved on from single-centre development to prospective evaluations, randomized trials, and national or programmatic rollouts that consistently demonstrate positive clinical impact on detection rates, workflow throughput, or both (\u003cspan additionalcitationids=\"CR13 CR14\" citationid=\"CR15\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e31\u003c/span\u003e). At the same time, a large number of studies are still retrospective or computational (train/validation/test on public or curated datasets), boasting high performance metrics that have not been externally validated or assessed in operational settings (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e21\u003c/span\u003e). In general, while several DL systems have been shown to produce compelling diagnostic performance in controlled and some real-world implementations, the evidence base is variable across modalities and geographies.\u003c/p\u003e\u003cp\u003e\u003cem\u003eInterpretation: innovation maturity across domains.\u003c/em\u003e\u003c/p\u003e\u003cp\u003eThe\u0026ensp;collective findings indicate a spectrum of maturity. Breast imaging has the most mature programmatic evidence: there are several large, prospective, population-based or real-world demonstration studies indicating that artificial intelligence (AI) can augment or replace a reader without loss of cancer detection and with potential gains in positive predictive value and efficiency (\u003cspan additionalcitationids=\"CR14\" citationid=\"CR16\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Gastroenterology CADe for\u0026ensp;polyp/adenoma detection is supported by randomized and multicenter trials indicating consistent improvements in adenoma and polyp detection rate, and a corresponding decrease in miss rates (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e20\u003c/span\u003e). A multitude of large-scale CXR TB screening solutions have matured to operational large-scale deployments, reporting high AUCs with measurable reductions in workload\u0026ensp;at an adapted operating point in screening workflows (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e31\u003c/span\u003e). In contrast, many dermatology, single-task radiology, ECG, and some prognostic applications are still largely in the development/retrospective validation stage: strong internal results are common, but independent external validation, prospective implementation studies, and long-term outcomes data are scarce or\u0026ensp;non-existent (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e25\u003c/span\u003e). Hence, the level of maturity of innovation is closely correlated with (a) the availability of large, programmatic datasets and (b) the existence of pragmatic prospective evaluations.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003eClinical implications\u003c/h2\u003e\u003cp\u003eOn current evidence, DL approaches are poised to enhance and support clinical services in specific, constrained tasks for which prospective and rollout data are available. These are (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) screening mammography workflows: as an independent reader, as a safety net, or triage in organized screening programs to substantially increase and faclitate cancer detection while reducing human workload under controlled deployment protocols (\u003cspan additionalcitationids=\"CR14\" citationid=\"CR16\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e15\u003c/span\u003e); (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2\u003c/span\u003e) population or programmatic CXR TB screening where high-throughput implementations have demonstrated safe triage thresholds and very high sensitivity with the potential for substantial automation (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e31\u003c/span\u003e); and (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e3\u003c/span\u003e) CADe for GI endoscopy for which RCT evidence supports routine use to reduce adenoma miss rates and enhance adenoma detection measures in multiple centres (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e20\u003c/span\u003e). For other tasks (dermoscopy, single-centre radiology algorithms, ECG arrhythmia classifiers, echocardiography prognostic models), DL may be suitable for decision-support or augmentation in research-savvy or pilot clinical environments but should not be considered for widespread unguided replacement of clinician judgment until comprehensive external validations, prospective effectiveness studies, and implementation activities have been conducted (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e25\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cem\u003eResearch and policy recommendations.\u003c/em\u003e\u003c/p\u003e\u003cp\u003eTo rapidly facilitate safe and equitable translation, the field should give priority to: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) multicenter, pragmatic randomized trials and prospective implementation studies that assess patient-relevant and cost-effectiveness outcomes rather than algorithmic metric performance alone; (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2\u003c/span\u003e) standardized validation hierarchies that mandate independent external validation cohorts, temporal\u0026ensp;validation, and the reporting of calibration and clinically relevant operating points; (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e3\u003c/span\u003e) prespecified subgroup analyses by key demographic and technical covariates (age, sex, race/ethnicity, device/vendor) and public dissemination of such\u0026ensp;findings; (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e) standardized postdeployment surveillance pipelines (continuous performance monitoring, drift detection, threshold recalibration) with publicly accessible run\u0026ensp;logs of model updates; (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e5\u003c/span\u003e) human-factors research to assess the impact of explainability tools (eg, saliency maps) on clinician decision-making and workflow; and (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e6\u003c/span\u003e) well-defined regulatory pathways and conflict-of-interest disclosure policies that facilitate independent audit of commercial systems in clinical environments. Where feasible, funders and journals should mandate data sharing or, at a minimum, standardized metadata descriptors\u0026ensp;to allow for replication and pooled analyses.\u003c/p\u003e\u003cp\u003e\u003cem\u003eStrengths and limitations of this scoping review.\u003c/em\u003e\u003c/p\u003e\u003cp\u003eThe strengths of this review include the detailed extraction of study design, validation method, implementation environment, and clinical outcome from 24 recent DL diagnostic studies, with an emphasis on the translation stage and real-world implementation. Additionally, it integrates common methodological and ethical issues that have a direct bearing on decision-making regarding implementation. Limitations are the scoping design (a quantitative meta-analysis was not conducted), potential publication and language bias, inclusion of diverse types of studies (clinical trials, prospective audits, computational studies) that impair comparability, and dependence on reporting within primary studies (which occasionally lacked complete metrics disclosures or subgroup data). Finally, there is some indication that new deployment studies may shift the evidence balance, given the rapidly evolving nature of the literature; ongoing updates and living evidence syntheses will be critical .\u003c/p\u003e\u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eDeep learning has been shown to reach clinically useful thresholds for selected diagnostic scenarios, including routine screening mammography, programmatic CXR TB screening, and colonoscopy CADe; in these areas, pragmatic large-scale evidence supports cautious adoption with operational safeguards. That said, for\u0026ensp;several other interesting applications, far more methodological bolstering, more extensive external validation, prospective outcome trials, equity-centered subgroup analyses, and mandated post-deployment surveillance are warranted before recommending widespread, unsupervised clinical adoption.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003cp\u003eNot applicable. This study did not involve human participants, human data, or tissue.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003cp\u003eNot applicable. No person\u0026rsquo;s data in any form is presented in this manuscript.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003cp\u003eAll authors certify that they have no affiliations with or involvement in any organization or entity with any financial interest (such as honoraria, educational grants, participation in speakers\u0026rsquo; bureaus, membership, employment, consultancies, stock ownership, or other equity interest; and expert testimony or patent-licensing arrangements), or non-financial interest (such as personal or professional relationships, affiliations, knowledge or beliefs) in the subject matter or materials discussed in this manuscript.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eGU conceived the study, designed the methodology, conducted the literature search, performed data extraction and synthesis, and drafted the manuscript. TO contributed to data verification, interpretation of findings, and drafting and critical revision of the final manuscript. Both authors reviewed and approved the final version for submission. GU is the guarantor of the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTopol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44\u0026ndash;56. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-018-0300-7\u003c/span\u003e\u003cspan address=\"10.1038/s41591-018-0300-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2019 Jan 7. PMID: 30617339.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEsteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115\u0026ndash;118. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nature21056\u003c/span\u003e\u003cspan address=\"10.1038/nature21056\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2017 Jan 25. Erratum in: Nature. 2017;546(7660):686. doi: 10.1038/nature22985. PMID: 28117445; PMCID: PMC8382232.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1:18. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-018-0029-1\u003c/span\u003e\u003cspan address=\"10.1038/s41746-018-0029-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTopol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44\u0026ndash;56. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-018-0300-7\u003c/span\u003e\u003cspan address=\"10.1038/s41591-018-0300-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2019 Jan 7. PMID: 30617339.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEsteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542(7639):115\u0026ndash;118. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nature21056\u003c/span\u003e\u003cspan address=\"10.1038/nature21056\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2017 Jan 25. Erratum in: Nature. 2017;546(7660):686. doi: 10.1038/nature22985. PMID: 28117445; PMCID: PMC8382232.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med. 2018;1:18. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-018-0029-1\u003c/span\u003e\u003cspan address=\"10.1038/s41746-018-0029-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAggarwal R, Sounderajah V, Martin G, Ting DSW, Karthikesalingam A, King D, Ashrafian H, Darzi A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. NPJ Digit Med. 2021;4(1):65. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-021-00438-z\u003c/span\u003e\u003cspan address=\"10.1038/s41746-021-00438-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 33828217; PMCID: PMC8027892.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYu AC, Mohajer B, Eng J. External Validation of Deep Learning Algorithms for Radiologic Diagnosis: A Systematic Review. Radiol Artif Intell. 2022;4(3):e210064. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/ryai.210064\u003c/span\u003e\u003cspan address=\"10.1148/ryai.210064\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 35652114; PMCID: PMC9152694.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOakden-Rayner L, Dunnmon J, Carneiro G, R\u0026eacute; C. Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging. Proc ACM Conf Health Inference Learn (2020). 2020;2020:151\u0026ndash;159. doi: 10.1145/3368555.3384468. PMID: 33196064; PMCID: PMC7665161.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHasanzadeh F, Josephson CB, Waters G, Adedinsewo D, Azizi Z, White JA. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit Med. 2025;8(1):154. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-025-01503-7\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-01503-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40069303; PMCID: PMC11897215.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFood US, Administration D. Proposed regulatory framework for modifications to artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD): discussion paper and request for feedback. FDA; 2019. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.fda.gov/media/122535/download\u003c/span\u003e\u003cspan address=\"https://www.fda.gov/media/122535/download\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSounderajah V, Ashrafian H, Golub RM, Shetty S, De Fauw J, Hooft L, STARD-AI Steering Committee, et al. Developing a reporting guideline for artificial intelligence-centred diagnostic test accuracy studies: the STARD-AI protocol. BMJ Open. 2021;11(6):e047709. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmjopen-2020-047709\u003c/span\u003e\u003cspan address=\"10.1136/bmjopen-2020-047709\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 34183345; PMCID: PMC8240576.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHe M, Li Z, Liu C, Shi D, Tan Z. Deployment of Artificial Intelligence in Real-World Practice: Opportunity and Challenge. Asia Pac J Ophthalmol (Phila). 2020 Jul-Aug;9(4):299\u0026ndash;307. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/APO.0000000000000301\u003c/span\u003e\u003cspan address=\"10.1097/APO.0000000000000301\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 32694344.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTricco AC, Lillie E, Zarin W, O\u0026rsquo;Brien KK, Colquhoun H, Levac D, Moher D, Peters MDJ, Horsley T, Weeks L et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med. 2018;169(7):467\u0026ndash;473. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.acpjournals.org\u003c/span\u003e\u003cspan address=\"https://www.acpjournals.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e/doi/10.7326/M18-0850\u003c/span\u003e\u003cspan address=\"/doi/10.7326/M18-0850\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZheng T, Lin F, Li X, Chu T, Gao J, Zhang S, Li Z, Gu Y, Wang S, Zhao F, Ma H, Xie H, Xu C, Zhang H, Mao N. Deep learning-enabled fully automated pipeline system for segmentation and classification of single-mass breast lesions using contrast-enhanced mammography: a prospective, multicentre study. EClinicalMedicine. 2023;58:101913. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eclinm.2023.101913\u003c/span\u003e\u003cspan address=\"10.1016/j.eclinm.2023.101913\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 36969336; PMCID: PMC10034267.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNg AY, Oberije CJG, Ambr\u0026oacute;zay \u0026Eacute;, Szab\u0026oacute; E, Serfőző O, Karpati E, Fox G, Glocker B, Morris EA, Forrai G, Kecskemethy PD. Prospective implementation of AI-assisted screen reading to improve early detection of breast cancer. Nat Med. 2023;29(12):3044\u0026ndash;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-023-02625-9\u003c/span\u003e\u003cspan address=\"10.1038/s41591-023-02625-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2023 Nov 16. PMID: 37973948; PMCID: PMC10719086.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDembrower K, Crippa A, Col\u0026oacute;n E, Eklund M, Strand F, ScreenTrustCAD. Trial Consortium. Artificial intelligence for breast cancer detection in screening mammography in Sweden: a prospective, population-based, paired-reader, non-inferiority study. Lancet Digit Health. 2023;5(10):e703-e711. doi: 10.1016/S2589-7500(23)00153-X. Epub 2023 Sep 8. Erratum in: Lancet Digit Health. 2023;5(10):e646. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S2589-7500(23)00181-4\u003c/span\u003e\u003cspan address=\"10.1016/S2589-7500(23)00181-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 37690911.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEisemann N, Bunk S, Mukama T, Baltus H, Elsner SA, Gomille T, Hecht G, Heywang-K\u0026ouml;brunner S, Rathmann R, Siegmann-Luz K, T\u0026ouml;llner T, Vomweg TW, Leibig C, Katalinic A. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat Med. 2025;31(3):917\u0026ndash;24. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-024-03408-6\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-03408-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2025 Jan 7. PMID: 39775040; PMCID: PMC11922743.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSalim M, Liu Y, Sorkhei M, Ntoula D, Foukakis T, Fredriksson I, Wang Y, Eklund M, Azizpour H, Smith K, Strand F. AI-based selection of individuals for supplemental MRI in population-based breast cancer screening: the randomized ScreenTrustMRI trial. Nat Med. 2024;30(9):2623\u0026ndash;30. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-024-03093-5\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-03093-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2024 Jul 8. PMID: 38977914; PMCID: PMC11405258.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGlissen Brown JR, Mansour NM, Wang P, Chuchuca MA, Minchenberg SB, Chandnani M, Liu L, Gross SA, Sengupta N, Berzin TM. Deep Learning Computer-aided Polyp Detection Reduces Adenoma Miss Rate: A United States Multi-center Randomized Tandem Colonoscopy Study (CADeT-CS Trial). Clin Gastroenterol Hepatol. 2022;20(7):1499\u0026ndash;e15074. Epub 2021 Sep 14. PMID: 34530161.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLachter J, Schlachter SC, Plowman RS, Goldenberg R, Raz Y, Rabani N et al. Novel artificial intelligence\u0026ndash;enabled deep learning system to enhance adenoma detection: a prospective randomized controlled study. iGIE. 2023;2(1):52\u0026ndash;8. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.igiejournal.org/article/S2949-7086(23)00015-8/fulltext\u003c/span\u003e\u003cspan address=\"https://www.igiejournal.org/article/S2949-7086(23)00015-8/fulltext\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePark DK, Kim EJ, Im JP, Lim H, Lim YJ, Byeon JS, Kim KO, Chung JW, Kim YJ. A prospective multicenter randomized controlled trial on artificial intelligence-assisted colonoscopy for enhanced polyp detection. Sci Rep. 2024;14(1):25453. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-024-77079-1\u003c/span\u003e\u003cspan address=\"10.1038/s41598-024-77079-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 39455850; PMCID: PMC11512038.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLagstr\u0026ouml;m RMB, Br\u0026auml;uner KB, Bielik J, Rosen AW, Crone JG, G\u0026ouml;genur I, Bulut M. Improvement in adenoma detection rate by artificial intelligence-assisted colonoscopy: Multicenter quasi-randomized controlled trial. Endosc Int Open. 2025;13:a25215169. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1055/a-2521-5169\u003c/span\u003e\u003cspan address=\"10.1055/a-2521-5169\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40018072; PMCID: PMC11866038.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDong L, He W, Zhang R, Ge Z, Wang YX, Zhou J, et al. Artificial Intelligence for Screening of Multiple Retinal and Optic Nerve Diseases. JAMA Netw Open. 2022;5(5):e229960. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2022.9960\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2022.9960\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 35503220; PMCID: PMC9066285.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBrant A, Singh P, Yin X, Yang L, Nayar J, Jeji D et al. Performance of a Deep Learning Diabetic Retinopathy Algorithm in India. JAMA Netw Open. 2025;8(3):e250984. doi: 10.1001/jamanetworkopen.2025.0984. Erratum in: JAMA Netw Open. 2025;8(4):e2511258. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2025.11258\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2025.11258\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40105843; PMCID: PMC11923701.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDai L, Sheng B, Chen T, Wu Q, Liu R, Cai C, et al. A deep learning system for predicting time to progression of diabetic retinopathy. Nat Med. 2024;30(2):584\u0026ndash;94. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-023-02702-z\u003c/span\u003e\u003cspan address=\"10.1038/s41591-023-02702-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub 2024 Jan 4. PMID: 38177850; PMCID: PMC10878973.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMarchetti MA, Cowen EA, Kurtansky NR, Weber J, Dauscher M, DeFazio J, Deng L, Dusza SW, Haliasos H, Halpern AC, Hosein S, Nazir ZH, Marghoob AA, Quigley EA, Salvador T, Rotemberg VM. Prospective validation of dermoscopy-based open-source artificial intelligence for melanoma diagnosis (PROVE-AI study). NPJ Digit Med. 2023;6(1):127. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-023-00872-1\u003c/span\u003e\u003cspan address=\"10.1038/s41746-023-00872-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 37438476; PMCID: PMC10338483.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWinkler JK, Blum A, Kommoss K, Enk A, Toberer F, Rosenberger A, Haenssle HA. Assessment of Diagnostic Performance of Dermatologists Cooperating With a Convolutional Neural Network in a Prospective Clinical Study: Human With Machine. JAMA Dermatol. 2023;159(6):621\u0026ndash;7. 2023.0905. PMID: 37133847; PMCID: PMC10157508.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbdullah M, Sadia Afrin, Mridha MF, Sultan Alfarhood, Che D, Safran M. Explainable deep learning approaches for high precision early melanoma detection using dermoscopic images. Sci Rep. 2025;15(1).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTrost B, Rodrigues L, Ong C, Dezellus A, Goldberg YH, Bouchat M et al. Artificial Intelligence Empowers Novice Users to Acquire Diagnostic-Quality Echocardiography. JACC: Advances. 2025;4(8):102005. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.sciencedirect.com/science/article/pii/S2772963X25004296\u003c/span\u003e\u003cspan address=\"https://www.sciencedirect.com/science/article/pii/S2772963X25004296\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLau ES, Di Achille P, Kopparapu K, Andrews CT, Singh P, Reeder C et al. Deep Learning-Enabled Assessment of Left Heart Structure and Function Predicts Cardiovascular Outcomes. Journal of the American College of Cardiology. 2023;82(20):1936\u0026ndash;48. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/37940231/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/37940231/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAtwa AEM, Atlam ES, Ahmed A, Atwa MA, Abdelrahim EM, Siam AI. Diagnostics (Basel). 2025;15(15):1950. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/diagnostics15151950\u003c/span\u003e\u003cspan address=\"10.3390/diagnostics15151950\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40804914; PMCID: PMC12346745. Interpretable Deep Learning Models for Arrhythmia Classification Based on ECG Signals Using PTB-X Dataset.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBai X, Dong X, Li Y, Liu R, Zhang H. A hybrid deep learning network for automatic diagnosis of cardiac arrhythmia based on 12-lead ECG. Sci Rep. 2024;14(1):24441. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-024-75531-w\u003c/span\u003e\u003cspan address=\"10.1038/s41598-024-75531-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 39424921; PMCID: PMC11489693.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMunjal P, Mahrooqi AA, Rajan R, Jeremijenko A, Ahmad I, Akhtar MI, Pimentel MAF, Khan S. Population-scale cross-sectional observational study for AI-powered TB screening on one million CXRs. NPJ Digit Med. 2025;8(1):418. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-025-01832-7\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-01832-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 40634545; PMCID: PMC12241593.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSharma V, None Nillmani, Gupta S, Shukla KK. Deep learning models for tuberculosis detection and infected region visualization in chest X-ray images. Intell Med. 2023.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMujeeb Rahman KK, Zulaikha S, Dhafer B, Ahmed R. Advancing tuberculosis screening: A tailored CNN approach for accurate chest X-ray analysis and practical clinical integration. Intelligence-Based Med. 2025;11:100196.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eIbrahim H, Liu X, Rivera SC, et al. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines. Trials. 2021;22:11. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s13063-020-04951-6\u003c/span\u003e\u003cspan address=\"10.1186/s13063-020-04951-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1 is available in the Supplementary Files section.\u003c/p\u003e\n"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7876598/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7876598/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e\u003cp\u003eDeep learning (DL) based diagnostic systems potentially offer automated image/signal interpretation and\u0026ensp;workflow support across a wide range of clinical fields, but the evidence for clinical translation is mixed. We conducted a scoping review to identify validation methods, evidence for implementation, and common\u0026ensp;methodological/operational challenges in current DL-based diagnostic research.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e\u003cp\u003eBased on PRISMA-ScR\u0026ensp;guidelines, we searched key bibliographic databases for peer-reviewed articles (2020\u0026ndash;2025) describing DL models for diagnostic tasks producing quantitative results. Two reviewers independently screened the records\u0026ensp;and extracted the study characteristics into a standardized data extraction form (Author, year, country, domain, task, sample size, model type, comparator, metrics, validation method, setting, key findings, reported challenges). We rated each study according to its furthest advanced\u0026ensp;stage of translation (development, external validation, prospective testing, randomized trial, post-deployment). No formal risk of bias assessment was conducted.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003eTwenty-four studies met the inclusion criteria across radiology (breast imaging, chest x-ray), gastroenterology (colonoscopy CADe), ophthalmology (retinal screening/prognostics), dermatology (dermoscopy), and cardiology (echocardiography/ECG). Mapping of the translation stage was found: prospective clinical testing n\u0026thinsp;=\u0026thinsp;8 (33.3%), randomized trials n\u0026thinsp;=\u0026thinsp;5 (20.8%), post-deployment/real-world auditing n\u0026thinsp;=\u0026thinsp;4 (16.7%), external validation n\u0026thinsp;=\u0026thinsp;2 (8.3%), and development/retrospective studies n\u0026thinsp;=\u0026thinsp;5 (20.8%). High-performing examples from prospective or rollout studies included higher cancer detection in screening mammography, randomized evidence of higher adenoma detection with CADe, large-scale TB CXR\u0026ensp;screening at AUCs\u0026thinsp;\u0026gt;\u0026thinsp;0.98 with workload reductions up to ~\u0026thinsp;80%, and echocardiography automation comparable to expert metrics. Frequently cited issues were\u0026ensp;a lack of geographic/demographic diversity, spectrum bias, inconsistent external validation, underreporting of clinically relevant operating points and calibration, gaps in explainability, barriers to integration with workflow, and variable regulatory/COI transparency.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e\u003cp\u003eDL diagnostic platforms have attained clinical-utility evidence in many applications (screening mammography, colonoscopy CADe, programmatic CXR screening) where prospective\u0026ensp;trials or deployments are available. That said, safe widespread adoption would require standardized external validation, prospective outcome studies, and analyses of equity-focused subgroups, routine post-deployment monitoring, and transparent reporting of thresholds and governance\u003c/p\u003e","manuscriptTitle":"Deep Learning in Clinical Diagnostics: A Scoping Review of Innovations Shaping Future Healthcare Delivery","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-18 08:19:15","doi":"10.21203/rs.3.rs-7876598/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-01-14T08:33:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"265489875762171497653654894246214209173","date":"2026-01-08T09:48:45+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-11T08:22:43+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"235346066006959867525908126662945927709","date":"2025-11-30T09:22:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"42255607244788058537192670619650059554","date":"2025-11-24T08:44:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"107689890247503428758207399839518016880","date":"2025-11-24T08:40:58+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-11-05T08:14:48+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-10-23T10:55:29+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-10-18T01:30:13+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-10-18T01:30:01+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Informatics and Decision Making","date":"2025-10-16T10:37:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"5cc7e9dd-5dbc-4549-acaf-875adb3f2e67","owner":[],"postedDate":"November 18th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-11-18T08:19:15+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-18 08:19:15","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7876598","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7876598","identity":"rs-7876598","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00