AI-Powered Clinical Decision Support Systems in Disease Diagnosis, Treatment Planning, and Prognosis: A Systematic Review.

OA: gold CC-BY-NC-4.0
AI-generated summary by claude@2026-07, 2026-07-29

This systematic review analyzed 123 studies on AI-powered clinical decision support systems, finding strong diagnostic and prognostic accuracy but noting issues with interpretability, data bias, and integration challenges.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

BackgroundArtificial intelligence (AI) is transforming healthcare with applications that can surpass human performance in prevention, detection, and treatment. This systematic review aimed to collect and assess the impact and success of AI technologies across various healthcare domains.MethodsA systematic search of major databases (including PubMed, Scopus, and ISI) was conducted for articles published up to 2023. Keywords related to AI-driven disease detection, classification, and prognosis were used. Non-English articles or those with inaccessible full texts were excluded. Data was extracted by two researchers, and the quality of selected articles was evaluated based on the strengths and limitations stated by the authors.ResultsIn total, 123 articles were included. AI contributions were categorized into three areas. For disease detection (n=75), Coronavirus disease 2019 (COVID-19) was the most frequent topic (n=18), followed by oncology. Chest X-rays were the most common input (n=15). In disease classification (n=23), oncology (especially breast cancer) was the most researched field (n=7), primarily using breast imaging. For prediction and prevention (n=25), oncology was again the most studied category, with clinical and laboratory parameters being the most utilized input (n=12).ConclusionAI-driven clinical decision support systems (CDSS) exhibit strong diagnostic and prognostic accuracy in imaging and laboratory settings. However, many models function as "black boxes," which limits interpretability and clinician trust. Data bias and challenges in integrating AI tools into practice also persist. The findings suggest that future work should focus on explainable AI and rigorous real-world validation to safely implement these tools in healthcare.
Full text 20,347 characters · extracted from pmc-nxml · 6 sections · click to expand

Methods

Search Strategy and Study Screening Process: This systematic review was conducted in 2024. Iran University of Medical Sciences approved the study (IR.IUMS.REC.1401.711). To select appropriate and relevant studies, an extensive electronic search was conducted. For this purpose, PubMed, ISI, Cochrane, Scopus, Embase, Science Direct, and Elsevier databases were searched. Articles published until 2023 were searched. To select studies, by applying Mesh term strategy we used keywords such as “AI powered,” “AI powered,” “AI- powered,” “AI assisted,” “AI assisted,” “AI based,” “AI based,” “AI enabled,” “AI enabled,” “AI aided,” “AI aided,” “Machine learning powered,” “Machine learning assisted,” “Machine learning based,” “Machine learning enabled,” “Machine learning aided,” “Deep learning powered,” “Deep learning assisted,” “Deep learning based,” “Deep learning enabled,” “Deep learning aided,” “Neural Network Computer,” “Computer based,” “Computer assisted,” “Computer enabled,” “Computer aided,” “Computer powered,” “System based,” “System assisted,” “System enabled,” “System aided,” “System powered,” “AI,” “Deep learning,” “Machine learning,” "Clinical Decision Support Systems," and "Clinical Decision Support”. Based on the main purpose of the study, the keywords of prevention, classification, and diagnosis (detection) were also used to search for studies. The types of included studies were intervention clinical trials, randomized controlled trials, case-control, prospective and retrospective cohort, and cross-sectional. Also, references to the selected articles were searched manually. For an extensive search, 3 researchers conducted the resource search process separately and eventually coordinated the selected studies. In the first search phase, 2024 articles were selected. Duplicated articles were detected by 1 researcher and supervised by a subsequent researcher using EndNote (X17) software, and 1250 articles were removed. The criteria used for duplication detection were similar in the titles, first author name, and the year of publication. Here, we focus on the open-access journals and those publicly available for the possibility of further investigation. A total of 126 articles were excluded due to full-text unavailability. The number of remaining articles after this process reached 648. Subsequently, the titles and abstracts of articles were evaluated based on inclusion criteria, and 243 articles met the inclusion criteria. Using a full-text review, 120 articles were excluded due to inappropriate content. Finally, 123 eligible studies were reviewed. The finding and screening flowchart was plotted using the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram tool ( 10 ) and reported in Figure 1 . We considered all studies with any study design as eligible for inclusion if they examined PICO as a tool ( Table 1 ) for developing a search strategy for identifying potentially relevant studies. We applied other restrictions in this review, such as studies related to the English language. Also, for quality assessment, we used evaluation criteria that are usual in artificial intelligence in medicine journals, including accuracy, precision, et cetera. We evaluate AI-based systems rather than the process of diagnosis. The articles whose full texts were not accessible were excluded. A standard checklist was developed for data collection and extraction. The design of the checklist was done under the supervision of the AI expert of the project. The designed checklist included the information of the extracted articles, such as the name of authors, year of publication, the country where the study took place, the type of AI model, the type of disease, the type of data processed, the AI performance measures and the limitations that mentioned in the manuscripts. To evaluate the quality of the selected articles, the strengths and limitations expressed in the articles by the authors were used as a proxy for the quality evaluation tool of the selected studies. Data extraction was done by 2 researchers of the study under the supervision of an AI expert. Any ambiguities and disagreements were listed and discussed in a session by these 2 researchers. If any problem remained after the discussion, it was investigated and resolved by the third person in the study. In the data-extracting process, the effort was to ensure that there was no missing data. The methodological quality and potential for bias of the selected studies, particularly those evaluating diagnostic accuracy, were assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. This tool evaluates studies across 4 key domains: patient selection, index test, reference standard, and flow and timing. Each domain was assessed for risk of bias, and the first 3 domains were also evaluated for concerns regarding applicability. Finding and screening flowchart

Results

This review included studies published until 2023. Through a systematic search of electronic databases and manual screening of references, 123 articles were identified. The findings of this review indicated that AI has made significant contributions to the field of medicine in various areas, including diagnosis, prognosis, and classification. Based on these findings, our included articles were classified into these three parts, presented in Table 2 to Table 4 . Our review identified a total of 75 studies that employed AI algorithms for the early detection and diagnosis of medical conditions. The details of each study, including the AI model used, performance measurement, data types, and study limitations (if available), are shown in Table 2 . These studies have covered a wide range of conditions, with respiratory infections, including COVID-19 disease, being the most commonly studied topic (in 18 out of the 75 studies). The second most researched field in this category was oncology, with a focus on gastrointestinal and skin cancers. AI was used in various applications for detection, ranging from the interpretation of medical images to the analysis of clinical and laboratory data. In this category, chest X-ray images were the most frequently used inputs (in 15 of the 75 studies), followed by CT scan images (in 9 of the 75 studies). It is worth noting that AI-based systems demonstrated high accuracy and efficiency in detecting abnormalities. In this category, out of 75 studies reviewed, 37 studies (49.3%) have employed “deep learning methods,” while 22 studies (29.3%) utilized “machine learning methods.” This demonstrates a significant reliance on deep learning algorithms for the accurate detection and diagnosis of medical conditions. In the domain of classification, our review identified 23 studies that employed AI techniques to categorize medical data into distinct classes. Table 3 presents the details of each study. These studies covered a wide range of medical conditions, with oncology (specifically breast cancer) being the most studied topic in 7 of 23 studies. Another widely researched area was cardiovascular diseases, which were addressed in 4 of 23 studies. AI algorithms demonstrated remarkable capabilities in accurately classifying patients based on various data, with breast images (ultrasound or mammograms) being the most used items. Out of 23 articles reviewed in this category, 11 studies (47.8%) have utilized “deep learning methods.” In addition, 4 studies (17.4%) have utilized “neural networks,” further highlighting the versatility and effectiveness of AI in classifying medical conditions. In the domain of prognosis and prevention, our review identified 25 studies that leveraged AI methodologies to predict the progression and outcomes of medical conditions, as shown in Table 4 . These studies covered a wide spectrum of diseases, ranging from heart failure to neurological disorders. Oncology again was the most studied field in this category. AI-based prognostic models exhibited impressive predictive performance, enabling clinicians to anticipate patient outcomes with greater accuracy and foresight. The category's most frequently used items were clinical and laboratory parameters, including PCR, used in 12 out of 25 articles. Furthermore, AI algorithms demonstrated the ability to identify prognostic factors that might otherwise go unnoticed, thereby facilitating early intervention and risk mitigation strategies. Out of 25 articles reviewed in this category, 14 studies (56%) have utilized “machine learning methods” for prognosis and prevention tasks. However, deep learning methods were only employed in 4 articles (16%). Furthermore, we evaluated research in the 3 classes based on sample size, AI models, and data type. Figure 2 shows the comparison from the aspect of sample size. As the result shows for the classification, the AI-based system required a larger sample size. Also, detection and diagnosis need a larger sample size than prognosis. Considering the AI models, we first investigated the AI-based model as presented in Figure 3 . The most used model is DL. Then, we analyzed them in each of the 3 classes. The most used model in detection-diagnosis and classification is the deep learning model, with about 49.3% and 47.8 %, respectively, and prognosis uses ML, with 56%. These results indicate that for detection and classification, deep learning is the most used model, while in prognosis, in which there is a need to find patterns among data, ML is the most used model that uses a smaller sample size than the deep learning model. We further evaluated the classes from the aspect of data type in AI-based systems. As the results in Table 2 to Table 4 indicate, in detection-diagnosis and classification, different modalities of images and signals are the 2 most used data types. More than 75% and 60% are images, and more than 15% and 26 % are signals in those 2 classes. Also, the most used data type in prognosis is laboratory data, with more than 53% usage. These results show that for detection and classification, images and signals carry proper information, while in prognosis, laboratory data analysis leads to extracting early signs of disease. The specific signaling questions used for the methodological quality and potential for bias of the selected studiesassessment, adapted from the standard QUADAS-2 checklist, are provided in Table 5 . The summary of the risk of bias and applicability judgments for the diagnostic accuracy studies included in this review (as listed in Table 2 ) is presented in Figure 4 . (A) Detection and Diagnosis, (B) classification, and (C) prognosis AI-based models in studies. Neural networks (NN), convolutional neural networks (CNN), and deep learning (DL) stand for 3 types of network-based ML algorithms, and ML class stands for the rest of the ML algorithms. Methodology checklist: the QUADAS-2 tool for studies of diagnostic test accuracy (listed according to table 2)

Conclusion

AI-based methods with a variety of approaches can be used in different areas of medicine, such as automated triage systems or real-time imaging analysis, based on the type of data and the amount of available data. DL, as the most popular strategy, has multiple processing layers in the network. They are used in many medical applications with high accuracy. In medical diagnoses, classification, and prognoses, the focus has been on the use of AI methods, which increase the accuracy, and criteria for selecting methods that increase the speed of diagnoses, classification, and prognoses were not provided. This may be a suggestion for future investigation. Also, the efficiency of AI-based methods compared with manual ones is discussed in most studies, which indicates the willingness to use AI in the medical industry. Also, with the human in the loop in medicine and medical staff duty, all AI-based applications are assistants, and the final choice is made by the human expert who uses such AI-based aid. As a result, no work is focused on replacing clinical professionals with AI. This demonstrates the complementarity and aid of artificial intelligence in medical services. It is indicated that in the future, it will be explored in which medical services AI may be fully implemented, as well as how human presence will be involved. Robotic surgery is one relevant example in this subject. Robotic surgery is one relevant example in this subject. This study was approved by the Ethics Committee of Iran University of Medical Sciences (IR.IUMS.REC.1401.711).

Discussion

The information obtained from studies was analyzed from several aspects, including the applied AI models, the investigated disease, sample size, data type, and measurement criteria. Most studies used structures based on DL in the model section. The most used methods were deep learning, ML, CNN, hybrid fuzzy-based learning systems, classical neural networks (NN), decision support systems (DSS), Bayesian network (BN), and particle swarm optimization (PSO). ML algorithms consist of two classes of non-network-based ML algorithms, such as random forest, and network-based ML algorithms, such as NN. Here, for precise classification, we separate each type of network-based algorithm of NN, CNN, and DL, which are mostly used in papers, and consider the rest of the ML algorithms as ML class. From the point of view of diseases, the two top investigated diseases were COVID-19 and cancer. Also, in terms of data type, various medical data types such as images, signals, clinical data, and geographical data were used in studies. However, medical images were used as the data investigated in most of the studies. In the sample size section, the largest data volume included 96,303 CT image samples, and the smallest included 19 MRI image samples. It shows that the proper sample size depends on the applied model. Besides, it should be noted that this sample size variation might affect results. Moreover, on the measurement criteria, accuracy, specificity, F1-score, and area under the curve were among the most used criteria in studies, and the accuracy of the model was at the top. The highest accuracy value of the models is equal to 99.9% in the diagnosis of COVID-19 using deep learning and 328 data points related to chest radiography, and also in the classification of diabetic macular edema using deep learning and 84000 tomographic images. CNNs are particularly well‑suited to high‑dimensional imaging tasks because their convolutional filters and pooling layers efficiently learn and aggregate spatial features, such as edges, textures, and anatomical structures, across multiple scales. In contrast, RNNs (and LSTM variants) excel at modeling temporal dependencies in sequential data, such as time‑series laboratory values or vital signs, by maintaining an internal state that captures information across prior time steps. The nature of the input modality further dictates model choice and hyperparameter tuning: structured clinical data (eg, tabular lab results, demographics) often employ fully connected networks or tree‑based models with feature engineering and regularization parameters (eg, learning rate, tree depth) optimized for tabular distributions, whereas unstructured imaging data require convolutional architectures with appropriately chosen filter sizes, depths, and spatial dropout rates to balance representational power and overfitting risk. Finally, hybrid CNN‑RNN pipelines enable multimodal integration—first extracting spatial embeddings from images via CNNs, then modeling their temporal evolution or combining them with sequential clinical measurements in RNN layers—thereby capturing both spatial and temporal patterns for more comprehensive clinical predictions. A key strength of this study is its comprehensive and systematic approach, as it analyzed a wide range of studies from multiple reputable databases, ensuring a thorough review of AI applications in medical diagnosis, classification, and prognosis. Additionally, the study highlighted AI's significant contributions across different medical fields, particularly in disease detection and prediction, providing valuable insights for future research. However, a notable weakness is the lack of emphasis on the speed and efficiency of AI models, as the study primarily focuses on accuracy without discussing the practical implications of AI adoption in clinical settings. Another limitation is its exclusion of non-English studies and studies with inaccessible full texts, which may introduce selection bias and limit the generalizability of its findings. Furthermore, while the paper categorizes AI applications effectively, it does not critically assess the challenges of AI implementation, such as ethical concerns, data biases, and real-world integration hurdles. Finally, considering the interpretability and data requirements, we should explain that DL models, compared to ML models, are less interpretable and need more data. However, using explainable artificial intelligence (XAI) techniques, this problem can be solved. Also, the mentioned limitations in the studies include lack of data, having small datasets that can be solved using data augmentation or federated learning, asymmetry of data, dependence on the data labeling process, incomplete or inaccurate data that lead to data bias, and the need for an expert's opinion in the data labeling process. Moreover, many high-performance DL models remain “black boxes,” limiting clinician trust; integrating XAI methods such as saliency mapping or SHAP can partially mitigate this but adds complexity. The reliance on English-language, publicly accessible datasets may also introduce geographic and demographic biases, and small sample sizes exacerbate overfitting; strategies like data augmentation and federated learning could improve generalizability. Practical integration into clinical workflows remains challenging due to workflow disruption, interoperability issues with electronic health records, and regulatory uncertainties, underscoring the need for clinician-AI–AI codesign. Finally, ethical and regulatory concerns, including patient privacy, accountability for AI-driven decisions, and the lack of standardized approval pathways—must be addressed to ensure safe and compliant deployment.

Introduction

Artificial Intelligence (AI) is a multidisciplinary study that aims to create a machine capable of perceiving data, inferring information, reaching intelligence, wisdom, cognition, and ultimately making decisions. AI aims to assist human beings in their decision-making by preparing a framework for processing all data at the same time, and presenting logical thinking and problem-solving ( 1 ). These machines are designed to handle vast, complex tasks without limitations in time or accuracy ( 2 ). According to the literature, AI is applied in a broad range of applications, technologies, and facilities, from software programs to robots. Like the other services and jobs, using AI in medical care is inevitable today ( 3 ). AI could compensate for some care delivery deficiencies, such as a lack of manpower and time-consuming tasks ( 2 ). Generally, AI is used in prevention, early detection of disease, and personalized and targeted therapy. Using a variety of computational tools, AI can process different types of data, such as laboratory, clinical, images, signals, et cetera, to find irregular patterns of disease, perform estimation, and prediction. The ability to work with multimodal data creates a holistic view for physicians to make quick and precise decisions. Studies show that AI may enable better disease prevention, diagnosis, and treatment. Among the main fields of disease that use AI tools, we can mention basic interventions in the fields of cancer, neurology, cardiology, and diabetes ( 4 - 6 ). For instance, many studies in radiologic diagnosis of different types of lung disease find the remarkable diagnostic value of various types of AI methods ( 7 - 9 ). However, there is a need for research to comprehensive study and address the effect of AI applications on healthcare. Therefore, here, we systematically reviewed the evidence to find the effect of different AI methods usage on the medical interventions classified as prevention, diagnosis, and treatment.

Coi Statement

The authors declare that they have no competing interests.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-10T06:11:17.106188+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-4.0