The Role of Agentic Artificial Intelligence in Healthcare: A Systematic Review | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article The Role of Agentic Artificial Intelligence in Healthcare: A Systematic Review Bernardo Gabriele Collaco, Syed Ali Haider, Srinivasagam Prabha, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7234499/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Mar, 2026 Read the published version in npj Digital Medicine → Version 1 posted 12 You are reading this latest preprint version Abstract Background/ Objectives: Agentic AI represents a promising evolution of AI technology applied to healthcare, with systems increasingly capable of operating autonomously to achieve defined clinical goals. However, the literature lacks conceptual clarity between “AI agents” and “agentic AI”, and few studies have rigorously explored their clinical applications. Therefore, this study aims to conduct a novel systematic review addressing this gap by examining agentic AI systems in healthcare settings, characterizing their applications, features, outcomes, and limitations, and clarifying the conceptual distinctions between AI agents and agentic AI systems using predefined and objective criteria. Methods A comprehensive search was conducted across PubMed, Embase, Cochrane, Scopus, and Google Scholar on April 6th, 2025. Studies were included if they involved AI systems in healthcare settings that demonstrated the following agentic features: autonomous operation, goal-directed behavior, and initiating action. Data on the clinical tasks achieved by the agents, key findings, features, and limitations were collected from the included studies. Screening and extraction followed PRISMA guidelines, with Risk of Bias assessed using ROBINS-I and Cochrane's Risk of Bias tools. Results Of 984 retrieved records, seven studies met the inclusion criteria, spanning domains such as emergency medicine, oncology, radiology, and rehabilitation. Multi-agent architecture was frequently used to decompose and coordinate complex workflows. Among the included studies, the AIs showed high accuracy in diagnosing cancer patients, conducting treatment plans, sending alerts, coaching messages, analysing image data, and adapting to challenging experimental scenarios. While demonstrating potential for improved efficiency, task accuracy, and patient engagement, significant limitations were noted: narrow task scope, lack of physical agency, limited clinical validation, and barriers to integration into real-world healthcare systems. Only one system had been deployed in a patient-facing trial setting. Conclusion The current literature suggests an emerging role and application of Agentic AI, holding promise with the potential to revolutionize diagnostics, triage, treatment planning, and patient management. However, real-world implementation and evaluations in the literature are limited. Future research must address critical validation, regulation, ethics, and clinical integration challenges to realize their full potential. Clear operational definitions and frameworks for evaluating agency are essential to support safe and effective deployment of these systems. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Agentic AI agentic systems AI agents autonomous systems healthcare Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. INTRODUCTION Over the past decade, artificial intelligence (AI) has undergone a profound and rapid transformation in medicine, demonstrating unprecedented performance on complex tasks, sometimes even approaching near-expert medical reasoning [ 1 ]. These advances have propelled AI beyond narrow, single-purpose tools [ 2 ], leading to the emergence of AI agents and agentic systems characterized by greater efficiency, autonomous operation, advanced reasoning, self-learning, and more dynamic interactions [ 3 ]. AI agents distinguish themselves from previous AIs by adapting autonomously with minimal human intervention, and integrating effectively with specialized models, often operating within sophisticated multi-agent frameworks and invoking external tools [ 3 , 4 ]. The ability of these agents to manage complex tasks collaboratively, offering a higher degree of operational autonomy and adaptability, is a hallmark of next-generation AI systems [ 5 ] (Fig. 1 ). Specifically, AI agents that exhibit the highest levels of decision-making, adaptability, and self-sufficiency are referred to as Agentic AIs [ 6 ]. However, the distinction between agentic AI and AI agents remains unclear in the current literature, as both refer to systems capable of autonomous actions and goal-directed behavior [ 6 , 7 ]. Additionally, although this breakthrough is paramount for transformative applications in healthcare, there is unfortunately a limited number of studies that rigorously explore their applications in the medical field. These promising agentic systems could autonomously analyze complex medical data, personalize treatment plans, and support clinical decision-making, ultimately improving efficiency and patient outcomes while easing clinician workload [ 8 ]. Therefore, this review aims to systematically assess the current applications of agentic AI systems in healthcare, identify prevailing definitions, features, limitations, and analyze the operational characteristics that distinguish them from other forms of AI. To our knowledge, this is the first systematic review investigating the specific applications of agentic AI systems in healthcare, while also proposing a conceptual clarification between AI agents and agentic AI. 2. METHODS To conduct the systematic review, the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) reporting guideline was followed in all stages of this study [9]. Additionally, a description of this study’s protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO) under protocol number CRD420251043917 [10]. A quantitative meta-analysis was not performed due to variations in study designs, population, and outcome measures. This heterogeneity would prevent a strong statistical analysis. Therefore, a meaningful narrative synthesis of the results of a systematic search was chosen to provide context-specific interpretation of the findings. a) Eligibility criteria Studies were included if they focused on Agentic AI systems applied in healthcare or clinical practice . The minimal core criteria used for considering an AI as an Agentic AI were Autonomous Operation, Goal-Directed Behavior (e.g., diagnosing diseases, optimizing processes), and Initiation of Actions/Tool Invocation (e.g., providing diagnostic outputs that trigger next steps, sending alerts) [11]. Peer-reviewed journal articles , including observational and interventional studies (randomized or non-randomized), published in English within the past 7 years (2018 onward) were considered for inclusion. This timeframe aligns with recent developments in AI agents and transformer-based architectures [12]. Studies were excluded if they described traditional rule-based AI or machine learning systems lacking agentic properties (e.g., static prediction, data processing tasks). Research unrelated to healthcare or medical domains (e.g., applications in finance, education, or general robotics) was also excluded unless a clear clinical connection could be established. In vitro/animal studies, abstracts without full text, inaccessible full texts, expert opinions, reviews, comments, brief correspondence, short communication, personal viewpoints, opinions, technical notes, non-abstracts, preclinical studies, letters to editors, editorials, duplicate records, publications not available in English, and published before 2018 were removed from consideration (this covers language exclusion and ensures no studies outside the date range are included implicitly). b) Databases searched The literature search was conducted across five digital bibliographic databases: PubMed, Embase, Cochrane, Scopus, and Google Scholar. c) Search strategy Two independent reviewers conducted a systematic search on April 6, 2025. The complete search strategy included a combination of the following keywords ("Artificial Intelligence" OR "Machine Learning" OR AI OR LLM OR "large language models" OR “GPT-4” OR “GPT-3” OR ChatGPT OR gemini OR Claude OR Bard OR Deepseek) AND ("Agentic AI" OR "autonomous AI" OR "AI agent" OR "AutoGPT" OR "Personal Assistant" OR "Research Agent" OR "multi-agent system" OR "workflow automation" OR "tool-using AI" OR "memory-enabled AI") AND ("health care" OR healthcare OR medical OR clinical OR medicine OR "clinical decision" OR diagnosis OR intervention OR management OR patient OR "patient education"). For PubMed, the Medical Subject Headings (MeSH) terms "Artificial Intelligence", "Clinical Medicine", and "Decision Making, Computer-Assisted" were included, and the correct adaptation for Embase was made. Moreover, for Google Scholar, more than 20,000 results were found, and only the initial 100 were chosen for review using Google Scholar’s relevance-based sorting, with the top 100 results selected to balance relevance, manageability, and reproducibility [13]. For Scopus, almost 80,000 results were found, so the search strategy needed to be narrowed to ("Artificial Intelligence" OR "Machine Learning" OR AI OR LLM OR "large language models" OR “GPT-4” OR ChatGPT) AND ("Agentic AI" OR "autonomous AI" OR "AI agent" OR "AutoGPT") AND ("health care" OR healthcare OR medical OR clinical OR medicine OR "clinical decision" OR diagnosis OR intervention OR management OR patient OR "patient education") . d) Study selection: Articles were imported into the EndNote software (version 21.3) [14], labeled, organized by each database, and deduplicated. In sequence, the papers were screened based on title and abstract, and the remaining studies were selected for complete text retrieval. The study selection, conducted by two independent reviewers, was based on the initial inclusion and exclusion criteria used to select the final articles for inclusion in the systematic review. No relevant conflicts were observed. All this process was executed in accordance with the PRISMA flow diagram (Figure 2). e) Data extraction Data from the included studies were extracted by one author and independently verified by a second author. All reviewers followed a standardized data extraction template developed in advance to maintain consistency. Any uncertainties during the process were discussed collaboratively, with a third reviewer available to resolve disagreements. Information gathered from each paper included: authors, publication year, country, study design, study domain, AI system name, main clinical task, key findings, features, and limitations. f) Risk of bias assessment Two authors independently conducted the bias assessment of the included studies. The Risk of Bias in Non-randomized Studies - of Interventions (ROBINS-I) tool was selected for its suitability in assessing bias in non-randomized studies of interventions, aligning with the heterogeneous experimental and observational designs included in this review [15]. For the single randomized controlled trial [16], Cochrane’s RoB-2 tool was used due to its widespread use and structured domain-based evaluation [17, 18]. Finally, a web tool was used to generate the figures used in this study [19]. These tools were chosen because they are validated, widely accepted in systematic review methodology, and offer detailed, domain-specific assessments tailored to our study designs. Any disagreements between reviewers were resolved through discussion and, when necessary, adjudicated by a third reviewer to ensure consensus. 3. RESULTS a) Study selection and baseline characteristics The initial search across five databases yielded 984 records. After removing duplicates and screening by title and abstract, 89 papers were retrieved for full-text assessment. Of these, 82 records were excluded based on our criteria: 53 did not meet the agentic AI inclusion criteria, 4 were not peer-reviewed, 3 lacked medical applications, 1 was an abstract only, 4 were reviews, 3 were Clinical Trial registries, and 14 were not fully available. Following this rigorous screening process, seven studies published between 2018 and 2025 were included in this systematic review (Figure 2). Study characteristics are reported in Table 1. Among the included studies, only one was a randomized controlled trial [16], with a total of 42 patients, 14 in each arm of the trial, evaluating three different approaches for lifestyle intervention among cancer survivors. The remaining articles were experimental and did not involve patient recruitment or intervention. They comprised simulation-based experiments [20, 21], a feasibility study employing a retrospective design [22], computational studies [23, 24], and a diagnostic accuracy study to evaluate exosome-based detection of hepatocellular carcinoma [25]. The USA [16, 24] and China [22, 23, 25] represented most of the countries where the studies occurred. Articles spanned various domains, primarily focusing on radiology and oncology. This limited number of high-quality clinical trials suggests that research on agentic AI in healthcare is still nascent, mainly relying on experimental and computational designs. Consequently, while demonstrating potential, the current evidence base is preliminary and requires more robust clinical validation. The heavy reliance on simulation-based and computational studies, with only a single small-scale randomized controlled trial, indicates that agentic AI applications in healthcare are predominantly in early-stage development. This highlights an urgent need for more rigorous, patient-centered research, particularly interventional studies, to establish clinical efficacy and safety. Agentic AI systems demonstrated autonomy in task execution and action initiation, requiring minimal human input to achieve specific goals. Their workflow environment results varied, ranging from non-significant effects [16] to excellent outcomes, including 94.1% accuracy in differentiating cancer from controls [25], 87% accuracy in adjusting gameplay [21], and achieving state-of-the-art performance in their practical scenarios [24]. This highlights that while the autonomous capabilities of agentic AI are promising, their effectiveness is highly dependent on the application and context, indicating a need for targeted development and rigorous validation to ensure consistent and significant clinical benefit. The observed architectural diversity, spanning single and multi-agent frameworks, alongside the pervasive integration of Large Language Models (LLMs), highlights a strategic effort to enhance the sophistication and utility of agentic AI in complex healthcare tasks. The included agents demonstrated single [16, 21, 24, 25] and multi-agent frameworks [20, 22, 23]. The named agent TraumaTracker was designed to function as a personal medical assistant during trauma resuscitation using two cooperating Belief-Desire-Intention (BDI) agents, one dedicated to documentation (Tracker Agent) and another to real-time alert generation (Alert Generator Agent) [20]. In contrast, GPT-Plan relied on several agents (Dosimetrist, Physicist, TPS_proxy, Human_proxy) to refine radiotherapy treatment plans, matching (cervical cancer) or exceeding (lung cancer) the performance of human experts [22]. Similarly, MultiMedRes implemented an LLM-based learner agent that decomposed complex medical questions and interacted with domain-specific expert models to perform zero-shot diagnostic reasoning on chest X-ray data [24]. ProtChat also employed a multi-agent framework, comprising the User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents, within a GPT-4-integrated system to automate complex protein analysis tasks [23]. Finally, the state-of-the-art results achieved by these agents, including integrating deep learning and retrieval-augmented LLMs in ChatExosome for hepatocellular carcinoma (HCC) detection, further illustrate the expanding capabilities of agentic AI in oncology and radiology [25]. The emergence of these multi-agent frameworks marks a clear progression toward addressing increasingly complex clinical problems by distributing specialized functions among collaborative AI agents. This division of labor enables sophisticated coordination, often achieving performance comparable to or surpassing human experts. Most of these systems integrated LLMs or multimodal LLMs with domain-specific models to improve task specificity, enhance learning efficiency, and support dynamic interactions across inference, evaluation, and visualization processes [22-25]. The reliance on LLMs reflects a critical trend in advancing agentic AI’s ability to comprehend complex data, engage in interactive reasoning, and generate human-like responses, thereby broadening their potential for advanced diagnostic support, clinical decision-making, and seamless integration into diverse medical workflows (Figure 3). Except for TraumaTracker [20], all AIs exhibited a certain degree of adaptability, restricted to the context in which they were working, but none have demonstrated adaptive behavior over time. Furthermore, several other limitations were noted. These included concerns regarding a narrow scope, bias, interpretability, dependence on the quality of training data, and challenges in real-world deployment, as six of the seven included studies have not yet been implemented in real-world settings [20-25]. Additionally, the complexity of developing reliable agentic behaviors, particularly in multi-agent systems, has led authors to emphasize the need for enhanced safety and ethical oversight. Table 1. Baseline characteristics of the included studies. USA = United States of America; RCT = randomized controlled trial; AI = artificial intelligence; VR = virtual reality; DVQA = Difference Visual Question Answering; GPT = Generative Pre-trained Transformer; PLLMs = protein large language models; RMSE = root mean square error; ROC = receiver operating characteristic; MAE = mean absolute error; PR = precision-recall; OARs = organs at risk; HCC = hepatocellular carcinoma; AFP = alpha fetoprotein; BDI = Belief-Desire-Intention; API = application programming interface; JSON = JavaScript Object Notation; TPS = treatment planning system; RAG = retrieval-augmented generation; FFT = feature fusion transformer; Q&A = question and answer. b) Quality assessment Figures 3 and 4 present the charts of included studies, evaluated using the RoB-2 tool for RCTs [17, 26] and the ROBINS-I tool for non-randomized studies [15], utilizing the Risk-Of-Bias VISualization (ROBVIS) tool [19]. Hassoon et al. (2021) had some concerns about bias due to participants’ awareness of the intervention, but did not raise any concerns about the other aspects. Across the six observational studies, the overall risk of bias ranged from moderate to critical, primarily due to the experimental or preclinical nature of the research. Croatti et al. (2019) and Mariselvam et al. (2023) were assessed as having a serious to critical risk of bias, respectively, due to the absence of control groups and reliance on simulations. However, they demonstrated promising systems with relevant applications and features that matched criteria. The other studies demonstrated an overall moderate risk of bias, primarily due to a lack of confounder adjustment and insufficient real-world clinical validation. 4. DISCUSSION This article represents the first systematic analysis of highly agentic AI systems in healthcare, evaluated against predefined objective criteria. While the included studies spanned diverse domains and clinical tasks, they revealed considerable similarities in their design and inherent limitations concerning the AI agents and the study methodologies. The following sections describe the overall concepts found in the literature, compare them with the key agentic features in the included studies, address the main cautions and constraints observed, outline the ethical considerations and limitations of this review, and provide future directions for agentic systems in healthcare. a) Overall Concepts AI systems have evolved significantly over time, with key differences emerging in how they function and interact with their environments, demonstrating increasing levels of autonomy and complexity. Traditional AIs are built to perform narrow, well-defined tasks, like classifying images or following preset rules, without the flexibility to adapt or generalize beyond their programming or instructions [6, 27]. Others are more than reactive, with the ability to generate text, images, or code by learning patterns from vast datasets. However, they still depend on user input and lack true initiative [28, 29]. Their rise was marked by the introduction of LLMs and the release of the Generative AI ChatGPT, developed by OpenAI in 2022 [28, 30]. Going further, AI agents are advanced proactive systems designed to pursue specific goals, make autonomous decisions, and interact dynamically with their surroundings, often within focused domains and with limited adaptability [31, 32]. On the other hand, the most sophisticated AI systems, known as agentic AIs, combine these abilities with a higher degree of independence, adaptability, and decision-making, allowing them to initiate actions, coordinate tasks, and respond intelligently to changing conditions with minimal human input. According to current literature, agentic systems can express several features; however, their classification varies among published studies. Their cornerstone and the two most present ones are [I] Autonomous Operation : the ability of an agent to operate without constant and direct human intervention, having greater control over actions executed [4] and [II] Goal-directed Behavior : the systems not only act based on their perception of the environment, but also monitor the effects of their actions and adjust them accordingly to achieve these goals [33]. To carry out their objectives, the systems often exhibit high levels of [III] adaptability to their environment, with continuous improvement over time, learning from past experiences, and optimizing workflows by leveraging advanced reasoning techniques, such as Chain-of-Thought, ReAct, and Tree-of-Thought [3, 11, 34]. To achieve this constant efficiency enhancement and meet these goals, these AIs require [IV] action initiation , decision-making, or tool invocation [3, 27, 31], which involves calling external APIs, databases, or software tools. Not only that, but a [V] multi-agent collaboration is frequently seen when a single agent cannot carry out complex tasks. AI agents can delegate tasks to each other or coordinate the workflow [35], facilitating the execution and delivery. Finally, [VI] long-term memory and [VII] retrieval-augmented/advanced reasoning complement the features that make agentic systems highly productive and operational [32, 36]. An overview of the various features that an agentic AI can express compared to other AI systems is presented in Figure 5. An important point to consider is that, although numerous papers [5, 6] and several blogs [37, 38] describing AI agents and the features of Agentic AIs were published in 2025, the underlying concepts are not novel. In fact, the first papers mentioning software agents and agency were published around the early 1990s [4, 35, 39], when researchers began to explore their capabilities. Therefore, it would be reasonable to think that the difference between AI agents and Agentic AIs would be well established in the literature. However, the exact definitions remain unclear in the current literature despite more than three decades since the mention of agents and agency [7, 11, 12, 32]. Authors frequently use these terms interchangeably, as they have the same meaning, which can lead to confusion regarding the underlying concepts. One possible explanation, also evident in papers included in this review, is that agentic features are not always present within an agent at the same level of complexity. In most applications, AIs exhibit varying degrees of these characteristics combined [5]. Alternatively, a different approach is to use the term "agentic AI systems" to describe systems that exhibit high levels of “agenticness” or agency, which is a property denoting autonomous and goal-directed behavior, rather than a fixed classification [11]. Generally speaking, the more features a system has, the more agency it will express [11, 40], and the more agentic it will be. As there is no consensus in the literature, and based on a reasonable approach [11], autonomous operation and goal-directed behavior were identified as the most relevant features to compose the criteria for AI agents to be considered agentic. Furthermore, the degree of adaptability in agentic AI systems varies widely across published studies, often lacking clear or consistent classifications. For this reason, action initiation (tool invocation) was the last component of the minimal required criteria of the agentic AIs to be included in this review. b) Autonomous Operation Current AI systems in healthcare exhibit increasingly autonomous behaviors, ranging from decision support to active patient engagement. TraumaTracker , published by Croatti et al. (2019), is a notable example of a highly agentic AI, capable of monitoring trauma workflows, generating over 430 reports in nine months, and autonomously triggering alerts for abnormal vital signs without continuous user input [20]. Similarly, Hassoon et al. (2021) developed SmartText . This AI-driven coaching intervention sent up to three personalized daily messages tailored to participants' schedules, physical characteristics, sensor data, preferences, and behavioral progress [16]. Although both systems were rule-based rather than learning-based, agentic AI evolved from rigid rule-driven designs to adaptive, personalized interventions operating in real time within two years. Mariselvam et al. (2023) marked another step forward by proposing a system capable of dynamically tailoring VR rehabilitation tasks through real-time skill assessment and autonomous strategy optimization, moving beyond behaviorally adaptive coaching [21]. In 2025, agentic systems advanced further, adopting multi-agent architectures that proactively invoke external models, integrate multimodal data, and self-correct through reflective reasoning [22, 23]. ProtChat by Huang et al. (2025) automated protein property predictions and protein−drug interactions without human intervention, while collaborating with expert models, LLMs (e.g., ChatGPT), multimodal reasoning, self-reflection, and API invocation [22-25]. Radiation oncology has emerged as a particularly promising field for agentic AI development, as demonstrated by Gu et al. (2025) [24], Wang et al. (2025) [22], and Yang et al. (2025) [25], who introduced experimental systems for diagnostic reasoning, treatment planning, and cancer detection. These developments reflect a significant shift: AI systems are becoming proactive collaborators capable of handling complex, adaptive, and multimodal tasks. However, despite patient data and steady advances, only one study involved real patients [16]. Most systems remain narrowly focused, rely on structured or simulated data, and lack validation in real-world clinical practice. This highlights that agentic AI systems are still in the early stages of development, with none yet achieving full autonomy in healthcare. c) Goal-directed Behavior A defining trait of agentic AI systems is their ability to pursue and adapt goals dynamically, moving beyond static, predefined responses toward strategies shaped by real-time context. Unlike traditional systems that execute fixed tasks, these agents monitor outcomes, evaluate progress, and adjust their paths to align with overarching objectives. For instance, although TraumaTracker and SmartText were programmed to deliver messages to users, these systems worked to optimize workflow in trauma centers and continuously serve as coaching AI, respectively [16, 20]. Moving forward, the reinforcement learning-based virtual AI assistant for children with Down syndrome continuously refined its therapeutic goals by interpreting user feedback and optimizing future interactions [21]. In radiotherapy, GPT-Plan exemplifies goal-directed reasoning by iteratively adjusting treatment plans through self-corrective loops, mimicking the adaptive deliberation of expert dosimetrists [22], and MultiMedRes extends this paradigm even further with proactive goal decomposition, breaking complex diagnostic challenges into adaptive subgoals that evolve as new information emerges [24]. Despite these advances, many current systems still display goal-directed and adaptive behavior only within narrow, task-specific boundaries, often constrained by rule-based logic. Expanding this capability to support temporal adaptability, learning, redefining goals across patient trajectories, and shifting clinical evidence remains a critical frontier. Such longitudinal goal management is particularly vital in healthcare, where long-term, personalized care depends on systems that can dynamically reprioritize objectives and strategies as conditions evolve. d) Action Initiation Another hallmark of agentic AI is the ability to independently trigger and coordinate actions, transforming them from passive decision aids into active participants in clinical workflows. Modern systems increasingly demonstrate this by autonomously invoking external tools, integrating resources, and orchestrating multi-agent collaboration. ChatExosome , for example, employs retrieval-augmented generation to autonomously query specialized databases and literature, generating evidence-based diagnostic insights without human prompting [25]. Similarly, ProtChat integrates GPT-4 with domain-specific protein language models, generating JSON outputs and seamlessly coordinating multiple analytic tools within a multi-agent framework to perform complex protein property predictions and interaction analyses [23]. The MultiMedRes framework demonstrates even more advanced orchestration, with its learner agent dynamically querying and integrating domain expert models to iteratively refine multimodal medical reasoning [24]. GPT-Plan demonstrates multi-agent collaboration in clinical workflows by autonomously initiating optimization processes and leveraging historical treatment data to enhance radiotherapy plans [22]. These capabilities illustrate a significant shift from passive decision support systems to active, self-directed agents that can mobilize knowledge and computational resources in real-time. In contrast, action-initiation mechanisms are absent in some highly autonomous diagnostic systems for conditions like diabetic retinopathy [41-44] or colon polyps [45], where generating a diagnosis or referral typically does not require invoking external tools or APIs. Nevertheless, this growing capacity for action-initiating autonomy raises essential concerns regarding control boundaries, auditability, and ensuring that all autonomous actions remain aligned with clinical objectives and safety standards. As these systems gain greater procedural independence, careful governance will be essential to mitigate risks and ensure their safe integration into healthcare environments. e) Caution and constraints Besides all the current applications mentioned, the promising future era in which Agentic AI is expected to contribute significantly to healthcare [34], it is not a given that they will deliver better results. For example, although SmartText exhibited better workflow automation, the results were inferior to those of MyCoach , the other intervention in the study [16]. This system was not an agentic AI because it relied on participants’ spoken intent to deliver coaching responses. This may suggest that providing patients with medical advice is more effective if they actively seek it, at least for lifestyle changes. Moreover, additional warnings in other fields need to be highlighted, raising concerns about the future of healthcare applications. Recent literature provides substantial evidence that agentic AI, despite its promise, has significant limitations in specific contexts. For tasks with predictable, fixed workflows, agentic AI's computational cost and latency make it unsuitable for real-time applications, while simpler automation frameworks offer greater efficiency [31]. In scenarios with low error tolerance, the susceptibility of agentic AI to hallucinations and incorrect reasoning poses unacceptable risks in critical decision-making contexts [38]. The "black box" nature of these systems, referred to as the non-explainability of their reasoning or internal workflow process [46], presents significant challenges. This lack of transparency can pose compliance risks and erode stakeholder trust in regulated industries, such as healthcare and finance [47]. Cybersecurity concerns are equally pressing because of vulnerabilities to manipulation, prompt injection attacks, and data privacy risks [38, 47]. Additionally, when predictable outcomes are essential, autonomous decision-making may be problematic for agentic AI applications, such as legal contracts, where precision and predictability are paramount [48]. These constraints collectively suggest that agentic AI should be deployed selectively, carefully considering task requirements and risk tolerance. Using agentic AI systems in routine, predictable workflows would be like using a sledgehammer to crack a nut; hence, they are best suited for complex, dynamic tasks that demand adaptability and autonomous decision-making. f) Ethical considerations Agentic AI is widely praised for its autonomy and potential to support clinicians in ethically complex decisions. Still, this very autonomy can also increase the risk of overreliance on systems that may not always be fully understood or properly regulated. A recent example is the Bioethics Artificial Intelligence Advisory (BAIA) framework introduced by Roy et al. (2025), which aims to assist with moral decision-making in healthcare. However, in such high-stakes situations, agentic systems must be more than technically capable, they must be built on high-quality, diverse data, routinely audited for bias and fairness, and designed to explain their reasoning in ways that healthcare professionals and families can clearly understand [49], avoiding to adopt a “black box” model at all costs [46]. Without these safeguards, critical questions will inevitably arise: Who is responsible if an autonomous AI causes harm? How should clinicians share authority with a system that acts independently? AI must be grounded in transparency and evidence-based design to address these concerns, ensuring accountability, patient safety, and trust in high-stakes clinical settings [50, 51]. g) Strengths and limitations This review identifies several key limitations in the included studies. Most were narrowly focused, used structured or simulated data, and lacked real-world clinical validation, with only one study involving actual patients [16]. Although highly agentic, the systems failed to demonstrate the ultimate level of agency according to current literature [3, 6, 11], presenting with no adaptive behavior over time and lacking long-term memory for other clinical settings. The few studies reflect the limited application of agentic AI in healthcare, and publication bias may have further restricted the evidence base. Additionally, inconsistent terminology across studies underscores the need for more precise definitions and standardized evaluation criteria [5]. However, this systematic review also provides essential strengths. It is the first comprehensive evaluation of agentic AI systems in healthcare, offering clear conceptual definitions, identifying their operational features, and synthesizing the current state of both single- and multi-agent architectures across diverse clinical domains. Besides, it aligns the analysis with practical constraints that must be addressed for safe and effective clinical adoption. h) Future Directions A multidisciplinary approach is essential to enhance system robustness, clinical relevance, and human-AI collaboration. Targeted use in high-demand areas, such as screening or chronic care, may offer the most immediate value. Still, broader integration will require infrastructure support, legal clarity, and ongoing research in this rapidly evolving field. The development of agentic AI systems in healthcare is still in its early stages, but the results of this review suggest a clear path toward more mature and impactful applications. Therefore, future work should focus on advancing these systems from proof-of-concept and simulation into real-world clinical environments, guided by structured benchmarks, multidisciplinary evaluation, and robust validation frameworks necessary to address all the gaps related to implementing AI agents. Hence, future development of agentic AI should focus on moving beyond experimental systems toward real-world, clinically integrated tools that function as intelligent collaborators rather than passive assistants [51]. These systems must incorporate transparency mechanisms and strengthen human-AI collaboration, ensuring they support rather than replace clinicians [52]. Additionally, the creation of holistic benchmarks is essential to evaluate technical performance, explainability, usability, and trust from a human-centered perspective, ensuring safety and real-world applicability [53]. Further randomized controlled trials with more sophisticated agent-based systems, extended follow-up periods, and larger, more diverse patient populations are necessary to support these advancements. 5. CONCLUSION In conclusion, although the development of agentic AI systems in healthcare is still in its early stages, they represent a transformative step in the evolution of healthcare technologies. These systems offer the potential to autonomously support clinical decision-making, optimize workflows, and enhance patient care through adaptive, goal-driven behaviors. Early applications show promise across diagnostics, treatment planning, and patient engagement; however, their real-world use remains limited and largely experimental, and standardized definitions are still needed. Ongoing research must focus on rigorous validation, ethical design, human-AI collaboration, and seamless integration into clinical environments to realize its full impact. As the field advances, agentic AI holds the potential to become a valuable partner in medicine, one that not only performs tasks but understands, learns, and acts in service of better health outcomes. To make this vision a reality, future efforts must also address practical challenges such as computational cost and latency, which can limit their suitability in time-sensitive clinical contexts. Enhancing processing efficiency, simplifying agent coordination, and leveraging technologies like edge computing will ensure these systems are effective and scalable in real-world healthcare settings. Declarations ACKNOWLEDGEMENTS: The figures were created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7. This study was funded by Mayo Clinic and the generosity of Eric and Wendy Schmidt; Dalio Philanthropies; and Gerstner Philanthropies. The funder played no role in study design, data collection, analysis and interpretation of data, or the writing of this manuscript. All authors declare no financial or non-financial competing interests. AUTHOR CONTRIBUTION: BC and AF helped with the conceptualization, and BC, AG, SH contributed to the methodology, including study screening, selection, and data extraction. Validation and formal analysis of the results was performed by BC. The original draft was prepared by BC and SH, with review and editing by SP, CG, AG, and AF. Supervision and project administration was the responsibility of AF. All authors read and approved of the final manuscript. DATA AVAILABILITY: The data supporting the findings of this systematic review are available within the article. CODE AVAILABILITY: Not applicable. References Pressman, S.M., et al., AI and Ethics: A Systematic Review of the Ethical Considerations of Large Language Model Use in Surgery Research . Healthcare (Basel), 2024. 12(8). Wilcox, A., M. Griffith, and O. Griffith, Looking Forward to AI and Medicine: Where Are We, and Where Are We Going? Mo Med, 2025. 122(1): p. 34–38. Karunanayake, N., Next-generation agentic AI for transforming healthcare . Informatics and Health, 2025. 2(2): p. 73–83. Wooldridge, M. and N.R. Jennings, Intelligent agents: theory and practice . The Knowledge Engineering Review, 1995. 10(2): p. 115–152. Hughes, L., et al., AI Agents and Agentic Systems: A Multi-Expert Analysis. Journal of Computer Information Systems: p. 1–29. Acharya, D.B., K. Kuppan, and B. Divya, Agentic AI: Autonomous Intelligence for Complex Goals - A Comprehensive Survey . IEEE Access, 2025. 13: p. 18912–18936. Ruan, J., et al., TPTU: large language model-based AI agents for task planning and tool usage. arXiv preprint arXiv:2308.03427, 2023. Maleki Varnosfaderani, S. and M. Forouzanfar, The Role of AI in Hospitals and Clinics: Transforming Healthcare in the 21st Century . Bioengineering (Basel), 2024. 11(4). Page, M.J., et al., The PRISMA 2020 statement: an updated guideline for reporting systematic reviews . BMJ, 2021. 372: p. n71. Collaco, B. THE ROLE OF AGENTIC AI IN HEALTHCARE: A SYSTEMATIC REVIEW. PROSPERO CRD420251043917 . 2025; Available from: https://www.crd.york.ac.uk/PROSPERO/view/CRD420251043917 . Shavit, Y., et al., Practices for governing agentic AI systems . Research Paper, OpenAI, 2023. Jheng, Y.C., et al., The era of artificial intelligence-based individualized telemedicine is coming . J Chin Med Assoc, 2020. 83(11): p. 981–983. Genovese, A., et al., Artificial intelligence in clinical settings: a systematic review of its role in language translation and interpretation . Ann Transl Med, 2024. 12(6): p. 117. Light, M., EndNote 1-2-3 Easy! Reference Management for the Professional, Second Edition , A. Agrawal, 2009, Springer, New York, USA, Price: €44.95, Soft Cover, 294 pages, ISBN: 978-0-387-95900-9, Website: www.springer.com. South African Journal of Botany, 2010. 76. Sterne, J.A., et al., ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions . BMJ, 2016. 355: p. i4919. Hassoon, A., et al., Randomized trial of two artificial intelligence coaching interventions to increase physical activity in cancer survivors . NPJ digital medicine, 2021. 4(1): p. 168. Higgins, J.P., et al., The Cochrane Collaboration's tool for assessing risk of bias in randomised trials . BMJ, 2011. 343: p. d5928. Cardoso, R., et al., An updated meta-analysis of novel oral anticoagulants versus vitamin K antagonists for uninterrupted anticoagulation in atrial fibrillation catheter ablation . Heart Rhythm, 2018. 15(1): p. 107–115. McGuinness, L.A. and J.P.T. Higgins, Risk-of-bias VISualization (robvis): An R package and Shiny web app for visualizing risk-of-bias assessments . Research Synthesis Methods, 2021. 12(1): p. 55–61. Croatti, A., et al., BDI personal medical assistant agents: The case of trauma tracking and alerting. Artificial Intelligence in Medicine, 2019. 96((Croatti A., [email protected] ; Montagna S., [email protected] ; Ricci A., [email protected] ) Computer Science and Engineering Department (DISI), University of Bologna, Campus of Cesena, Via dell'Università 50, Cesena, Italy): p. 187–197. Mariselvam, J., S. Rajendran, and Y. Alotaibi, Reinforcement learning-based AI assistant and VR play therapy game for children with Down syndrome bound to wheelchairs . AIMS Mathematics, 2023. 8(7): p. 16989–17011. Wang, Q., et al., A feasibility study of automating radiotherapy planning with large language model agents . Physics in medicine and biology, 2025. 70(7). Huang, H., et al., ProtChat: An AI Multi-Agent for Automated Protein Analysis Leveraging GPT-4 and Protein Language Model . Journal of chemical information and modeling, 2025. 65(1): p. 62–70. Gu, Z., et al., A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning . Advanced Intelligent Systems, 2025. Yang, Z., et al., ChatExosome: An Artificial Intelligence (AI) Agent Based on Deep Learning of Exosomes Spectroscopy for Hepatocellular Carcinoma (HCC) Diagnosis. Analytical chemistry, 2025. 97(8): p. 4643–4652. Hassoon, A., et al., Increasing Physical Activity Amongst Overweight and Obese Cancer Survivors Using an Alexa-Based Intelligent Agent for Patient Coaching: Protocol for the Physical Activity by Technology Help (PATH) Trial . JMIR Res Protoc, 2018. 7(2): p. e27. Ogbu, D., Agentic AI in Computer Vision Domain - Recent Advances and Prospects . International Journal of Research Publication and Reviews, 2024. 5(11): p. 5102–5120. Lee, H., The rise of ChatGPT: Exploring its potential in medical education . Anat Sci Educ, 2024. 17(5): p. 926–931. Hadi, M.U., et al., Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects . Authorea Preprints, 2023. 1: p. 1–26. Sanderson, K., GPT-4 is here: what scientists think . Nature, 2023. 615(7954): p. 773. Sapkota, R., K.I. Roumeliotis, and M. Karkee, AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenge. arXiv preprint arXiv:2505.10468, 2025. Schneider, J., Generative to Agentic AI: Survey, Conceptualization, and Challenges. arXiv preprint arXiv:2504.18875, 2025. Castelfranchi, C., Modelling social action for AI agents . Artificial Intelligence, 1998. 103(1): p. 157–182. Hosseini, S. and H. Seilani, The role of agentic AI in shaping a smart future: A systematic review . Array, 2025. 26. Fischer, K., et al. Sophisticated and distributed: The transportation domain . in Proceedings of 9th IEEE Conference on Artificial Intelligence for Applications . 1993. Liu, G., et al., Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception. arXiv preprint arXiv:2505.23275, 2025. Viswanathan, S.P.a.M. 5 Levels of agentic AI intelligence for enterprise use . Outshift 2025 05/26/2025]; Available from: https://outshift.cisco.com/blog/agentic-ai-intelligence-for-enterprise-use . The Six Levels of Agentic Behavior . 2025 05/24/25]; Available from: https://www.vellum.ai/blog/levels-of-agentic-behavior . Demazeau, Y. and J.P. Müller, Decentralized A.I . 1990: North Holland. Chan, A., et al., Harms from Increasingly Agentic Algorithmic Systems , in 2023 ACM Conference on Fairness Accountability and Transparency . 2023. p. 651–666. Abràmoff, M.D., et al., Mitigation of AI adoption bias through an improved autonomous AI system for diabetic retinal disease . NPJ digital medicine, 2024. 7(1): p. 369. Wolf, R.M., et al., Autonomous artificial intelligence increases screening and follow-up for diabetic retinopathy in youth: the ACCESS randomized control trial . Nature communications, 2024. 15(1): p. 421. Abràmoff, M.D., et al., Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices . NPJ digital medicine, 2018. 1: p. 39. Abramoff, M.D., et al., Autonomous artificial intelligence increases real-world specialist clinic productivity in a cluster-randomized trial . NPJ digital medicine, 2023. 6(1): p. 184. Djinbachian, R., et al., AUTONOMOUS ARTIFICIAL INTELLIGENCE VERSUS AI ASSISTED HUMAN OPTICAL DIAGNOSIS OF COLORECTAL POLYPS: a RANDOMIZED CONTROLLED TRIAL . Gastrointestinal endoscopy, 2024. 99(6): p. AB11. Marcus, E. and J. Teuwen, Artificial intelligence and explanation: How, why, and when to explain black boxes . Eur J Radiol, 2024. 173: p. 111393. Simoes, B.S. Agentic AI: A Promising Evolution, But Not Without Limits . 2025 05/26/25]; Available from: https://www.blueprintsys.com/blog/agentic-ai-a-promising-evolution-but-not-without-limits . Warmly. AI Agentic Workflows: Definitions, Use Cases & Software . 2025 05/26/25]; Available from: https://www.warmly.ai/p/blog/ai-agentic-workflows . Dutta Roy, T.P., Bioethics Artificial Intelligence Advisory (BAIA): An Agentic Artificial Intelligence (AI) Framework for Bioethical Clinical Decision Support . Cureus, 2025. 17(3): p. e80494. Papagni, G., et al., Artificial agents’ explainability to support trust: considerations on timing and context . Ai & Society, 2022. 38(2): p. 947–960. Papagni, G. and S. Koeszegi, A pragmatic approach to the intentional stance semantic, empirical and ethical considerations for the design of artificial agents . Minds and Machines, 2021. 31(4): p. 505–534. Heer, J., Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences, 2019. 116(6): p. 1844–1850. Gridach, M., et al., Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions. arXiv preprint arXiv:2503.08979, 2025. Tables Table 1. Baseline characteristics of included studies Study Country Study Type Domain Agentic AI Name Main Clinical Task Key Findings Features Limitations Croatti 2019 [20] Italy Experimental Emergency Medicine TraumaTracker Real-time trauma tracking and generation of clinical alerts during resuscitation Improved documentation quality, enabled real-time alerts aligned with trauma team workflows - Autonomous trauma event tracking; - Goal-directed behavior for optimizing workflow with multimodal interaction: tablet, smart glasses; - Real-time reports and alert generation - Short-term memory during each trauma case; - Multi-Agent collaboration: two cooperating BDI agents; - Narrow task scope: trauma only; - Limited domain; - Reliance on structured resuscitation scenarios; - Tool and integration constraints; - Limited clinical validation and generalizability; - Requires manual input for some tasks; - Over-alerting risk in specific rules; - Customization and interoperability improvements needed; - Not adaptive in a dynamic or machine learning sense; - No long-term memory; Hassoon 2021 [16] USA RCT Lifestyle Intervention SmartText Personalized physical activity coaching to increase daily step counts Had a modest, non-significant effect - Autonomous message generation; - Goal-directed behavior as a coaching system; - Sends three personalized text messages per day; - Adaptive behavior based on participant data; - Short-term memory by storing recent physical activity; - Single Agent system; - Narrow scope: increase footsteps; - Limited domain; - Short study duration: 4 weeks; - Small sample size: 42 patients; - Limited generalizability - Not adapted to complex clinical environments; - Limited adaptability: structured, rule-driven, and preprogrammed setting, not adaptable over time; - No long-term memory; Mariselvam 2023 [21] Saudi Arabia Experimental Pediatric rehabilitation Reinforcement Learning-Based Virtual AI Assistant Enhancing motor, cognitive, and social skills in children with Down syndrome using VR play therapy PPO + Actor-Critic performed best with 87% accuracy in adapting difficulty and supporting skill development; suitable for independent gameplay - Autonomous decision-making in a simulated setting; - Goal-directed behavior for skills improvement; - Action initiation by tilting the virtual board, moving the ball, and introducing obstacles; - Adaptive behavior by adjusting game difficulty, analyzing performance, and modifying strategies; - Short-term memory during training and gameplay; - Single Agent system; - Narrow scope: single VR therapy game; - Limited domain; - Reliance on sub-modules; - Lack of physical agency; - Limited clinical validation and generalizability: children who used wheelchairs with Down syndrome; - No real-world autonomy; - Limited adaptability: restricted to a single game environment, not adaptable over time; - No long-term memory; Gu 2025 [24] USA Experimental Radiology MultiMedRes DVQA between sequential chest X-rays Achieved state-of-the-art performance on DVQA tasks; GPT-4-based learner agent outperformed both fine-tuned and fully supervised models; effective in zero-shot settings - Autonomously determines when to stop querying; - Goal-directed behavior to address medical multimodal reasoning problems; - Initiation of external actions by launching API calls on its own; - Adaptive behavior by changing its reasoning strategy; - Short-term, task-specific memory; - Single LLM agent + expert models (APIs); - Retrieval-Augmented Reasoning; - Narrow Scope: DVQA for chest X-rays; - Limited domain; - Reliance on expert modules; - Lack of physical agency; - Limited clinical validation and generalizability; - No real-world autonomy; - Absence of patient-specific context; - Limited adaptability: task-specific and within constraints, not adaptable over time; - No long-term memory; Huang 2025 [23] China Experimental Biomedicine ProtChat Automation of complex protein analysis tasks using GPT-4 and PLLMs Accurately automated protein analysis tasks with strong performance across benchmark datasets, using standard evaluation metrics, such as Pearson, RMSE and MAE, ROC curve, and PR curve. - Autonomous task planning and execution via GPT-4; - Goal-directed behavior for protein understanding tasks; - Initiation of external actions: invokes PLLMs, runs functions, writes JSON results, and passes them to downstream agents - Adaptive behavior by handling multiple tasks; - Short-term, session-bound memory; - Multi-agent collaboration: User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents; - Narrow scope: protein analysis tasks; - Limited domain; - Reliance on sub-modules; - Lack of physical agency; - Limited clinical validation: Only tested on computational protein datasets; - Limited adaptability: task-specific and instruction-driven, not adaptable over time; - No long-term memory; Wang 2025 [22] China Experimental Radiation Oncology GPT-Plan Optimization and generation of radiotherapy treatment plans Matched or outperformed expert human planners and auto-planning systems in plan quality and efficiency, especially effective in OAR sparing and optimization iterations for lung and cervical cancer - Autonomous plan generation; - Goal-directed behavior for treatment optimization; - Initiates external actions/tool calls - Adaptive behavior by refining treatment plans; - Multi-agent collaboration: Dosimetrist, Physicist, TPS_proxy, Human_proxy; - Retrieval-Augmented Reasoning: Retriever tool; - Narrow scope: lung and cervical cancer; - Limited domain; - Reliance on sub-modules; - Lack of physical agency; - Small sample size: 17 patients; - Limited clinical validation and generalizability; - Limited adaptability: not across sessions, not adaptable over time; - No long-term memory; Yang 2025 [25] China Experimental Oncology ChatExosome Diagnose hepatocellular carcinoma from plasma exosome Raman spectroscopy ChatExosome achieved 94.1% accuracy in distinguishing HCC from controls, including 87.5% in AFP-negative HCC cases - Autonomous execution of multi-step reasoning; - Goal-directed loop and decision-making; - External tool invocation (FFT classifier, RAG retrieval) in response to file or text inputs; - Adaptive behavior by switching between predefined models - Short-term memory from a predefined external knowledge base; - Single Agent system - Retrieval-Augmented Reasoning - Partially narrow scope: demonstrated capacity of more than one type of cancer diagnosis; - Single domain; - Reliance on pretrained model; - Lack of physical agency; - Limited clinical validation and generalizability; - Limited adaptability: does not self-update; not adaptable over time; - No long-term memory; Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 14 Mar, 2026 Read the published version in npj Digital Medicine → Version 1 posted Editorial decision: Revision requested 13 Sep, 2025 Reviews received at journal 12 Sep, 2025 Reviews received at journal 10 Sep, 2025 Reviews received at journal 13 Aug, 2025 Reviewers agreed at journal 08 Aug, 2025 Reviewers agreed at journal 08 Aug, 2025 Reviewers agreed at journal 06 Aug, 2025 Reviewers agreed at journal 05 Aug, 2025 Reviewers invited by journal 04 Aug, 2025 Editor assigned by journal 04 Aug, 2025 Submission checks completed at journal 04 Aug, 2025 First submitted to journal 28 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7234499","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":496285786,"identity":"8605826a-e21b-45f0-b9a7-3abfb0c81d20","order_by":0,"name":"Bernardo Gabriele Collaco","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Bernardo","middleName":"Gabriele","lastName":"Collaco","suffix":""},{"id":496285787,"identity":"11053769-62fc-491f-9de5-d9f62fe31300","order_by":1,"name":"Syed Ali Haider","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Syed","middleName":"Ali","lastName":"Haider","suffix":""},{"id":496285788,"identity":"460f92da-138b-40d5-940e-e15b0c62cc81","order_by":2,"name":"Srinivasagam Prabha","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Srinivasagam","middleName":"","lastName":"Prabha","suffix":""},{"id":496285789,"identity":"f28a1556-0fc2-414b-b28b-69237ecba593","order_by":3,"name":"Cesar Abraham Gomez-Cabello","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Cesar","middleName":"Abraham","lastName":"Gomez-Cabello","suffix":""},{"id":496285790,"identity":"f8ccaafe-35d1-44f7-aa65-930cbb47bf83","order_by":4,"name":"Ariana Genovese","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Ariana","middleName":"","lastName":"Genovese","suffix":""},{"id":496285791,"identity":"b7dacc9e-4746-4bdb-b168-304cf2693bf4","order_by":5,"name":"Nadia G. Wood","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Nadia","middleName":"G.","lastName":"Wood","suffix":""},{"id":496285792,"identity":"1aa4b49a-8872-47fd-8115-538c8a137cf7","order_by":6,"name":"Sanjay P. Bagaria","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Sanjay","middleName":"P.","lastName":"Bagaria","suffix":""},{"id":496285793,"identity":"479f473e-4a8f-4b36-af79-63ba671e60b8","order_by":7,"name":"Narayanan Gopala","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Narayanan","middleName":"","lastName":"Gopala","suffix":""},{"id":496285795,"identity":"7a1774aa-dfab-4c66-b0ac-2f24c7d71fe0","order_by":8,"name":"Cui Tao","email":"","orcid":"","institution":"Mayo Clinic","correspondingAuthor":false,"prefix":"","firstName":"Cui","middleName":"","lastName":"Tao","suffix":""},{"id":496285798,"identity":"43ed9f28-c4f4-476f-a01c-f5c7463890ca","order_by":9,"name":"Antonio Jorge Forte","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYFCCAwwfQBQ/AwMbTIgNt2qIFsYZDAwGDJINxGthgGgxOECsFv7GM4YNH/f8kTO+kfzswYcKBnl+sQNsjyvwaJE4cMawccYzA2OzG2nmhjPOMBjOnJ3AbngGr1eOpT/mOWCQuO1Ggpk0bxtDgsHtBDagx3AD+QPHEpuBWuo3z0j/RpwWgwOHD4K0JBhI5BBpiyFQS+OMA8ZAb7wpk5xxRgLol8R2Q3xa5G4cbGz4cEBOnr89fZvEhwobeX7p5GMP8WkBBhmUIZAA5gIxI14NwIiByfMfwKNqFIyCUTAKRjQAAKoYUYeSZNNNAAAAAElFTkSuQmCC","orcid":"","institution":"Mayo Clinic","correspondingAuthor":true,"prefix":"","firstName":"Antonio","middleName":"Jorge","lastName":"Forte","suffix":""}],"badges":[],"createdAt":"2025-07-28 13:38:24","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7234499/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7234499/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41746-026-02517-5","type":"published","date":"2026-03-14T15:59:39+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":88507191,"identity":"db212ad6-5bb3-41f2-98c2-4351167afc92","added_by":"auto","created_at":"2025-08-07 07:36:59","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":296603,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eAgentic AI Workflow. Created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/17abc522d55864825ef5faf2.jpeg"},{"id":88507100,"identity":"35b629c3-d92e-4936-8f60-71c2e526b646","added_by":"auto","created_at":"2025-08-07 07:36:36","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":105673,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eThe PRISMA flow diagram of the study selection process.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/3eccb07a14ffc8ebbcc3fe3e.png"},{"id":88507057,"identity":"464bff47-76f1-435b-b9e1-5cf0e5e0d7f1","added_by":"auto","created_at":"2025-08-07 07:36:20","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":181293,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eCurrent applications of Agentic AI in healthcare.\u003c/em\u003e Created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/c03cf30dcc636d0b081c7927.jpeg"},{"id":88507140,"identity":"8ac60630-c153-435b-a496-34cc8a051ce3","added_by":"auto","created_at":"2025-08-07 07:36:44","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":78285,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eRisk of Bias Chart for the included RCT using the RoB-2 tool proposed by Cochrane. The study demonstrated some overall concerns about bias. This figure was generated using the publicly available ROBVIS tool.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/0be3bee1cd509827195e85d9.png"},{"id":88507062,"identity":"5995b0d4-cac4-4c9c-bcfe-eb012b96d846","added_by":"auto","created_at":"2025-08-07 07:36:22","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":196334,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eRisk of Bias Chart for the observational studies using the ROBINS-I tool proposed by Cochrane. Croatti et al. (2019) and Mariselvam et al. (2023) demonstrated an overall serious and critical risk of bias, respectively, whereas the other studies demonstrated an overall moderate risk of bias. This figure was generated using the publicly available ROBVIS tool.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/4c6b74f3512ab6be24973c1c.png"},{"id":88507175,"identity":"ccaa4d8f-b13d-46b5-8ae3-a131ce0c93cc","added_by":"auto","created_at":"2025-08-07 07:36:56","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":179581,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eA comparison between the key features of Traditional AI, Generative AI, AI Agents, and Agentic AI. \u003c/em\u003eCreated in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7\u003c/p\u003e","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/199c76bd1e339b99de2585ee.jpeg"},{"id":104739477,"identity":"f6b205b6-bbe0-4808-b08c-e48c6b2a42d0","added_by":"auto","created_at":"2026-03-16 16:07:26","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1738120,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7234499/v1/e632f0ac-35d6-43ad-83c1-a539b9e64535.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eThe Role of Agentic Artificial Intelligence in Healthcare: A Systematic Review\u003c/p\u003e","fulltext":[{"header":"1. INTRODUCTION","content":"\u003cp\u003eOver the past decade, artificial intelligence (AI) has undergone a profound and rapid transformation in medicine, demonstrating unprecedented performance on complex tasks, sometimes even approaching near-expert medical reasoning [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. These advances have propelled AI beyond narrow, single-purpose tools [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], leading to the emergence of AI agents and agentic systems characterized by greater efficiency, autonomous operation, advanced reasoning, self-learning, and more dynamic interactions [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eAI agents distinguish themselves from previous AIs by adapting autonomously with minimal human intervention, and integrating effectively with specialized models, often operating within sophisticated multi-agent frameworks and invoking external tools [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The ability of these agents to manage complex tasks collaboratively, offering a higher degree of operational autonomy and adaptability, is a hallmark of next-generation AI systems [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Specifically, AI agents that exhibit the highest levels of decision-making, adaptability, and self-sufficiency are referred to as Agentic AIs [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. However, the distinction between agentic AI and AI agents remains unclear in the current literature, as both refer to systems capable of autonomous actions and goal-directed behavior [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eAdditionally, although this breakthrough is paramount for transformative applications in healthcare, there is unfortunately a limited number of studies that rigorously explore their applications in the medical field. These promising agentic systems could autonomously analyze complex medical data, personalize treatment plans, and support clinical decision-making, ultimately improving efficiency and patient outcomes while easing clinician workload [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eTherefore, this review aims to systematically assess the current applications of agentic AI systems in healthcare, identify prevailing definitions, features, limitations, and analyze the operational characteristics that distinguish them from other forms of AI. To our knowledge, this is the first systematic review investigating the specific applications of agentic AI systems in healthcare, while also proposing a conceptual clarification between AI agents and agentic AI.\u003c/p\u003e"},{"header":"2. METHODS","content":"\u003cp\u003eTo conduct the systematic review, the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) reporting guideline was followed in all stages of this study [9]. Additionally, a description of this study\u0026rsquo;s protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO) under protocol number CRD420251043917 [10]. A quantitative meta-analysis was not performed due to variations in study designs, population, and outcome measures. This heterogeneity would prevent a strong statistical analysis. Therefore, a meaningful narrative synthesis of the results of a systematic search was chosen to provide context-specific interpretation of the findings.\u003c/p\u003e\n\u003cp\u003ea) \u0026nbsp; \u003cem\u003eEligibility criteria\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eStudies were included if they focused on \u003cstrong\u003eAgentic AI systems\u003c/strong\u003e applied in \u003cstrong\u003ehealthcare or clinical practice\u003c/strong\u003e. The minimal core criteria used for considering an AI as an Agentic AI were \u003cstrong\u003eAutonomous Operation, Goal-Directed Behavior\u003c/strong\u003e (e.g., diagnosing diseases, optimizing processes), \u003cstrong\u003eand Initiation of Actions/Tool Invocation\u003c/strong\u003e (e.g., providing diagnostic outputs that trigger next steps, sending alerts) [11]. Peer-reviewed \u003cstrong\u003ejournal articles\u003c/strong\u003e, including observational and interventional studies (randomized or non-randomized), published in \u003cstrong\u003eEnglish\u003c/strong\u003e within the \u003cstrong\u003epast 7 years\u003c/strong\u003e (2018 onward) were considered for inclusion. This timeframe aligns with recent developments in AI agents and transformer-based architectures [12].\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;Studies were excluded if they described traditional rule-based AI or machine learning systems lacking agentic properties (e.g., static prediction, data processing tasks). Research unrelated to healthcare or medical domains (e.g., applications in finance, education, or general robotics) was also excluded unless a clear clinical connection could be established. In vitro/animal studies, abstracts without full text, inaccessible full texts, expert opinions, reviews, comments, brief correspondence, short communication, personal viewpoints, opinions, technical notes, non-abstracts, preclinical studies, letters to editors, editorials, duplicate records, publications not available in English, and published before 2018 were removed from consideration \u0026nbsp;(this covers language exclusion and ensures no studies outside the date range are included implicitly).\u003c/p\u003e\n\u003cp\u003eb) \u0026nbsp; \u003cem\u003eDatabases searched\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe literature search was conducted across five digital bibliographic databases: PubMed, Embase, Cochrane, Scopus, and Google Scholar.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ec) \u0026nbsp; \u003cem\u003eSearch strategy\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eTwo independent reviewers conducted a systematic search on April 6, 2025. The complete search strategy included a combination of the following keywords (\u0026quot;Artificial Intelligence\u0026quot; OR \u0026quot;Machine Learning\u0026quot; OR AI OR LLM OR \u0026quot;large language models\u0026quot; OR \u0026ldquo;GPT-4\u0026rdquo; OR \u0026ldquo;GPT-3\u0026rdquo; OR ChatGPT OR gemini OR Claude OR Bard OR Deepseek) AND (\u0026quot;Agentic AI\u0026quot; OR \u0026quot;autonomous AI\u0026quot; OR \u0026quot;AI agent\u0026quot; OR \u0026quot;AutoGPT\u0026quot; OR \u0026quot;Personal Assistant\u0026quot; OR \u0026quot;Research Agent\u0026quot; OR \u0026quot;multi-agent system\u0026quot; OR \u0026quot;workflow automation\u0026quot; OR \u0026quot;tool-using AI\u0026quot; OR \u0026quot;memory-enabled AI\u0026quot;) \u0026nbsp;AND (\u0026quot;health care\u0026quot; OR healthcare OR medical OR clinical OR medicine OR \u0026quot;clinical decision\u0026quot; OR diagnosis OR intervention OR management OR patient OR \u0026quot;patient education\u0026quot;).\u003c/p\u003e\n\u003cp\u003eFor PubMed, the Medical Subject Headings (MeSH) terms \u0026quot;Artificial Intelligence\u0026quot;, \u0026quot;Clinical Medicine\u0026quot;, and \u0026quot;Decision Making, Computer-Assisted\u0026quot; were included, and the correct adaptation for Embase was made. Moreover, for Google Scholar, more than 20,000 results were found, and only the initial 100 were chosen for review using Google Scholar\u0026rsquo;s relevance-based sorting, with the top 100 results selected to balance relevance, manageability, and reproducibility [13]. For Scopus, almost 80,000 results were found, so the search strategy needed to be narrowed to \u003cem\u003e(\u0026quot;Artificial Intelligence\u0026quot; OR \u0026quot;Machine Learning\u0026quot; OR AI OR LLM OR \u0026quot;large language models\u0026quot; OR \u0026ldquo;GPT-4\u0026rdquo; OR ChatGPT) AND (\u0026quot;Agentic AI\u0026quot; OR \u0026quot;autonomous AI\u0026quot; OR \u0026quot;AI agent\u0026quot; OR \u0026quot;AutoGPT\u0026quot;) \u0026nbsp;AND (\u0026quot;health care\u0026quot; OR healthcare OR medical OR clinical OR medicine OR \u0026quot;clinical decision\u0026quot; OR diagnosis OR intervention OR management OR patient OR \u0026quot;patient education\u0026quot;)\u003c/em\u003e. \u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003ed) \u0026nbsp; \u003cem\u003eStudy selection:\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eArticles were imported into the EndNote software (version 21.3) [14], labeled, organized by each database, and deduplicated. In sequence, the papers were screened based on title and abstract, and the remaining studies were selected for complete text retrieval. The study selection, conducted by two independent reviewers, was based on the initial inclusion and exclusion criteria used to select the final articles for inclusion in the systematic review. No relevant conflicts were observed. All this process was executed in accordance with the PRISMA flow diagram (Figure 2).\u003c/p\u003e\n\u003cp\u003ee) \u0026nbsp; \u003cem\u003eData extraction\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eData from the included studies were extracted by one author and independently verified by a second author. All reviewers followed a standardized data extraction template developed in advance to maintain consistency. Any uncertainties during the process were discussed collaboratively, with a third reviewer available to resolve disagreements. Information gathered from each paper included: authors, publication year, country, study design, study domain, AI system name, main clinical task, key findings, features, and limitations.\u003c/p\u003e\n\u003cp\u003ef) \u0026nbsp; \u0026nbsp;\u003cem\u003eRisk of bias assessment\u0026nbsp;\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eTwo authors independently conducted the bias assessment of the included studies. The Risk of Bias in Non-randomized Studies - of Interventions (ROBINS-I) tool was selected for its suitability in assessing bias in non-randomized studies of interventions, aligning with the heterogeneous experimental and observational designs included in this review [15]. For the single randomized controlled trial [16], Cochrane\u0026rsquo;s RoB-2 tool was used due to its widespread use and structured domain-based evaluation [17, 18]. Finally, a web tool was used to generate the figures used in this study [19]. These tools were chosen because they are validated, widely accepted in systematic review methodology, and offer detailed, domain-specific assessments tailored to our study designs. Any disagreements between reviewers were resolved through discussion and, when necessary, adjudicated by a third reviewer to ensure consensus.\u003c/p\u003e"},{"header":"3. RESULTS","content":"\u003cp\u003e\u003cem\u003ea) \u0026nbsp; Study selection and baseline characteristics\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThe initial search across five databases yielded 984 records. After removing duplicates and screening by title and abstract, 89 papers were retrieved for full-text assessment. Of these, 82 records were excluded based on our criteria: 53 did not meet the agentic AI inclusion criteria, 4 were not peer-reviewed, 3 lacked medical applications, 1 was an abstract only, 4 were reviews, 3 were Clinical Trial registries, and 14 were not fully available. Following this rigorous screening process, seven studies published between 2018 and 2025 were included in this systematic review (Figure 2). Study characteristics are reported in Table 1.\u003c/p\u003e\n\u003cp\u003eAmong the included studies, only one was a randomized controlled trial [16], with a total of 42 patients, 14 in each arm of the trial, evaluating three different approaches for lifestyle intervention among cancer survivors. The remaining articles were experimental and did not involve patient recruitment or intervention. They comprised simulation-based experiments [20, 21], a feasibility study employing a retrospective design [22], computational studies [23, 24], and a diagnostic accuracy study to evaluate exosome-based detection of hepatocellular carcinoma [25]. The USA [16, 24] and China [22, 23, 25] represented most of the countries where the studies occurred. Articles spanned various domains, primarily focusing on radiology and oncology. This limited number of high-quality clinical trials suggests that research on agentic AI in healthcare is still nascent, mainly relying on experimental and computational designs. Consequently, while demonstrating potential, the current evidence base is preliminary and requires more robust clinical validation. The heavy reliance on simulation-based and computational studies, with only a single small-scale randomized controlled trial, indicates that agentic AI applications in healthcare are predominantly in early-stage development. This highlights an urgent need for more rigorous, patient-centered research, particularly interventional studies, to establish clinical efficacy and safety.\u003c/p\u003e\n\u003cp\u003eAgentic AI systems demonstrated autonomy in task execution and action initiation, requiring minimal human input to achieve specific goals. Their workflow environment results varied, ranging from non-significant effects [16] to excellent outcomes, including 94.1% accuracy in differentiating cancer from controls [25], 87% accuracy in adjusting gameplay [21], and achieving state-of-the-art performance in their practical scenarios [24]. This highlights that while the autonomous capabilities of agentic AI are promising, their effectiveness is highly dependent on the application and context, indicating a need for targeted development and rigorous validation to ensure consistent and significant clinical benefit.\u003c/p\u003e\n\u003cp\u003eThe observed architectural diversity, spanning single and multi-agent frameworks, alongside the pervasive integration of Large Language Models (LLMs), highlights a strategic effort to enhance the sophistication and utility of agentic AI in complex healthcare tasks. The included agents demonstrated single [16, 21, 24, 25] and multi-agent frameworks [20, 22, 23]. The named agent \u003cem\u003eTraumaTracker\u003c/em\u003e was designed to function as a personal medical assistant during trauma resuscitation using two cooperating Belief-Desire-Intention (BDI) agents, one dedicated to documentation (Tracker Agent) and another to real-time alert generation (Alert Generator Agent) [20]. In contrast, \u003cem\u003eGPT-Plan\u003c/em\u003e relied on several agents (Dosimetrist, Physicist, TPS_proxy, Human_proxy) to refine radiotherapy treatment plans, matching (cervical cancer) or exceeding (lung cancer) the performance of human experts [22]. Similarly, \u003cem\u003eMultiMedRes\u003c/em\u003e implemented an LLM-based learner agent that decomposed complex medical questions and interacted with domain-specific expert models to perform zero-shot diagnostic reasoning on chest X-ray data [24]. \u003cem\u003eProtChat\u003c/em\u003e also employed a multi-agent framework, comprising the User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents, within a GPT-4-integrated system to automate complex protein analysis tasks [23]. Finally, the state-of-the-art results achieved by these agents, including integrating deep learning and retrieval-augmented LLMs in \u003cem\u003eChatExosome\u0026nbsp;\u003c/em\u003efor hepatocellular carcinoma (HCC) detection, further illustrate the expanding capabilities of agentic AI in oncology and radiology [25].\u003c/p\u003e\n\u003cp\u003eThe emergence of these multi-agent frameworks marks a clear progression toward addressing increasingly complex clinical problems by distributing specialized functions among collaborative AI agents. This division of labor enables sophisticated coordination, often achieving performance comparable to or surpassing human experts. Most of these systems integrated LLMs or multimodal LLMs with domain-specific models to improve task specificity, enhance learning efficiency, and support dynamic interactions across inference, evaluation, and visualization processes [22-25]. The reliance on LLMs reflects a critical trend in advancing agentic AI\u0026rsquo;s ability to comprehend complex data, engage in interactive reasoning, and generate human-like responses, thereby broadening their potential for advanced diagnostic support, clinical decision-making, and seamless integration into diverse medical workflows (Figure 3).\u003c/p\u003e\n\u003cp\u003eExcept for \u003cem\u003eTraumaTracker\u0026nbsp;\u003c/em\u003e[20], all AIs exhibited a certain degree of adaptability, restricted to the context in which they were working, but none have demonstrated adaptive behavior over time. Furthermore, several other limitations were noted. These included concerns regarding a narrow scope, bias, interpretability, dependence on the quality of training data, and challenges in real-world deployment, as six of the seven included studies have not yet been implemented in real-world settings [20-25]. Additionally, the complexity of developing reliable agentic behaviors, particularly in multi-agent systems, has led authors to emphasize the need for enhanced safety and ethical oversight.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eTable 1. Baseline characteristics of the included studies. USA = United States of America; RCT = randomized controlled trial; AI = artificial intelligence; VR = virtual reality; DVQA = Difference Visual Question Answering; GPT = Generative Pre-trained Transformer; PLLMs = protein large language models; RMSE = root mean square error; ROC = receiver operating characteristic; MAE = mean absolute error; PR = precision-recall; OARs = organs at risk; HCC = hepatocellular carcinoma; AFP = alpha fetoprotein; BDI = Belief-Desire-Intention; API = application programming interface; JSON = JavaScript Object Notation; TPS = treatment planning system; RAG = retrieval-augmented generation; FFT = feature fusion transformer; Q\u0026amp;A = question and answer.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eb) \u0026nbsp; Quality assessment\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eFigures 3 and 4 present the charts of included studies, evaluated using the RoB-2 tool for RCTs [17, 26] and the ROBINS-I tool for non-randomized studies [15], utilizing the Risk-Of-Bias VISualization (ROBVIS) tool [19]. Hassoon et al. (2021) had some concerns about bias due to participants\u0026rsquo; awareness of the intervention, but did not raise any concerns about the other aspects. Across the six observational studies, the overall risk of bias ranged from moderate to critical, primarily due to the experimental or preclinical nature of the research. Croatti et al. (2019) and Mariselvam et al. (2023) were assessed as having a serious to critical risk of bias, respectively, due to the absence of control groups and reliance on simulations. However, they demonstrated promising systems with relevant applications and features that matched criteria. The other studies demonstrated an overall moderate risk of bias, primarily due to a lack of confounder adjustment and insufficient real-world clinical validation.\u003c/p\u003e"},{"header":"4. DISCUSSION","content":"\u003cp\u003eThis article represents the first systematic analysis of highly agentic AI systems in healthcare, evaluated against predefined objective criteria. While the included studies spanned diverse domains and clinical tasks, they revealed considerable similarities in their design and inherent limitations concerning the AI agents and the study methodologies. The following sections describe the overall concepts found in the literature, compare them with the key agentic features in the included studies, address the main cautions and constraints observed, outline the ethical considerations and limitations of this review, and provide future directions for agentic systems in healthcare.\u003c/p\u003e\n\u003cp\u003ea) \u0026nbsp; \u003cem\u003eOverall Concepts\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eAI systems have evolved significantly over time, with key differences emerging in how they function and interact with their environments, demonstrating increasing levels of autonomy and complexity. Traditional AIs are built to perform narrow, well-defined tasks, like classifying images or following preset rules, without the flexibility to adapt or generalize beyond their programming or instructions [6, 27]. Others are more than reactive, with the ability to generate text, images, or code by learning patterns from vast datasets. However, they still depend on user input and lack true initiative [28, 29]. Their rise was marked by the introduction of LLMs and the release of the Generative AI ChatGPT, developed by OpenAI in 2022 [28, 30]. Going further, AI agents are advanced proactive systems designed to pursue specific goals, make autonomous decisions, and interact dynamically with their surroundings, often within focused domains and with limited adaptability [31, 32]. On the other hand, the most sophisticated AI systems, known as agentic AIs, combine these abilities with a higher degree of independence, adaptability, and decision-making, allowing them to initiate actions, coordinate tasks, and respond intelligently to changing conditions with minimal human input.\u003c/p\u003e\n\u003cp\u003eAccording to current literature, agentic systems can express several features; however, their classification varies among published studies. Their cornerstone and the two most present ones are\u003cstrong\u003e\u0026nbsp;[I] Autonomous Operation\u003c/strong\u003e: the ability of an agent to operate without constant and direct human intervention, having greater control over actions executed [4] and \u003cstrong\u003e[II] Goal-directed Behavior\u003c/strong\u003e: the systems not only act based on their perception of the environment, but also monitor the effects of their actions and adjust them accordingly to achieve these goals [33]. To carry out their objectives, the systems often exhibit high levels of \u003cstrong\u003e[III]\u003c/strong\u003e \u003cstrong\u003eadaptability\u003c/strong\u003e to their environment, with continuous improvement over time, learning from past experiences, and optimizing workflows by leveraging advanced reasoning techniques, such as Chain-of-Thought, ReAct, and Tree-of-Thought [3, 11, 34]. To achieve this constant efficiency enhancement and meet these goals, these AIs require \u003cstrong\u003e[IV] action initiation\u003c/strong\u003e, decision-making, or tool invocation [3, 27, 31], which involves calling external APIs, databases, or software tools. Not only that, but a \u003cstrong\u003e[V]\u003c/strong\u003e \u003cstrong\u003emulti-agent collaboration\u003c/strong\u003e is frequently seen when a single agent cannot carry out complex tasks. AI agents can delegate tasks to each other or coordinate the workflow [35], facilitating the execution and delivery. Finally, \u003cstrong\u003e[VI] long-term memory\u003c/strong\u003e and \u003cstrong\u003e[VII] retrieval-augmented/advanced reasoning\u003c/strong\u003e complement the features that make agentic systems highly productive and operational [32, 36]. An overview of the various features that an agentic AI can express compared to other AI systems is presented in Figure 5.\u003c/p\u003e\n\u003cp\u003eAn important point to consider is that, although numerous papers [5, 6] and several blogs [37, 38] describing AI agents and the features of Agentic AIs were published in 2025, the underlying concepts are not novel. In fact, the first papers mentioning software agents and agency were published around the early 1990s [4, 35, 39], when researchers began to explore their capabilities.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTherefore, it would be reasonable to think that the difference between AI agents and Agentic AIs would be well established in the literature. However, the exact definitions remain unclear in the current literature despite more than three decades since the mention of agents and agency [7, 11, 12, 32]. Authors frequently use these terms interchangeably, as they have the same meaning, which can lead to confusion regarding the underlying concepts. One possible explanation, also evident in papers included in this review, is that agentic features are not always present within an agent at the same level of complexity. In most applications, AIs exhibit varying degrees of these characteristics combined [5]. Alternatively, a different approach is to use the term \u0026quot;agentic AI systems\u0026quot; to describe systems that exhibit high levels of \u0026ldquo;agenticness\u0026rdquo; or agency, which is a property denoting autonomous and goal-directed behavior, rather than a fixed classification [11]. Generally speaking, the more features a system has, the more agency it will express [11, 40], and the more agentic it will be.\u003c/p\u003e\n\u003cp\u003eAs there is no consensus in the literature, and based on a reasonable approach [11], \u003cstrong\u003eautonomous operation\u003c/strong\u003e and \u003cstrong\u003egoal-directed behavior\u003c/strong\u003e were identified as the most relevant features to compose the criteria for AI agents to be considered agentic. Furthermore, the degree of adaptability in agentic AI systems varies widely across published studies, often lacking clear or consistent classifications. For this reason, \u003cstrong\u003eaction initiation\u003c/strong\u003e (tool invocation) was the last component of the minimal required criteria of the agentic AIs to be included in this review.\u003c/p\u003e\n\u003cp\u003eb) \u0026nbsp; \u003cem\u003eAutonomous Operation\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eCurrent AI systems in healthcare exhibit increasingly autonomous behaviors, ranging from decision support to active patient engagement. \u003cem\u003eTraumaTracker\u003c/em\u003e, published by Croatti et al. (2019), is a notable example of a highly agentic AI, capable of monitoring trauma workflows, generating over 430 reports in nine months, and autonomously triggering alerts for abnormal vital signs without continuous user input [20]. Similarly, Hassoon et al. (2021) developed\u003cem\u003e\u0026nbsp;SmartText\u003c/em\u003e. This AI-driven coaching intervention sent up to three personalized daily messages tailored to participants\u0026apos; schedules, physical characteristics, sensor data, preferences, and behavioral progress [16].\u003c/p\u003e\n\u003cp\u003eAlthough both systems were rule-based rather than learning-based, agentic AI evolved from rigid rule-driven designs to adaptive, personalized interventions operating in real time within two years. Mariselvam et al. (2023) marked another step forward by proposing a system capable of dynamically tailoring VR rehabilitation tasks through real-time skill assessment and autonomous strategy optimization, moving beyond behaviorally adaptive coaching [21].\u003c/p\u003e\n\u003cp\u003eIn 2025, agentic systems advanced further, adopting multi-agent architectures that proactively invoke external models, integrate multimodal data, and self-correct through reflective reasoning [22, 23]. \u003cem\u003eProtChat\u003c/em\u003e by Huang et al. (2025) automated protein property predictions and protein\u0026minus;drug interactions without human intervention, while collaborating with expert models, LLMs (e.g., ChatGPT), multimodal reasoning, self-reflection, and API invocation\u003cem\u003e\u0026nbsp;\u003c/em\u003e[22-25].\u003c/p\u003e\n\u003cp\u003eRadiation oncology has emerged as a particularly promising field for agentic AI development, as demonstrated by Gu et al. (2025) [24], Wang et al. (2025) [22], and Yang et al. (2025) [25], who introduced experimental systems for diagnostic reasoning, treatment planning, and cancer detection. These developments reflect a significant shift: AI systems are becoming proactive collaborators capable of handling complex, adaptive, and multimodal tasks. However, despite patient data and steady advances, only one study involved real patients [16]. Most systems remain narrowly focused, rely on structured or simulated data, and lack validation in real-world clinical practice. This highlights that agentic AI systems are still in the early stages of development, with none yet achieving full autonomy in healthcare.\u003c/p\u003e\n\u003cp\u003ec) \u0026nbsp; \u003cem\u003eGoal-directed Behavior\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eA defining trait of agentic AI systems is their ability to pursue and adapt goals dynamically, moving beyond static, predefined responses toward strategies shaped by real-time context. Unlike traditional systems that execute fixed tasks, these agents monitor outcomes, evaluate progress, and adjust their paths to align with overarching objectives. For instance, although TraumaTracker and \u003cem\u003eSmartText\u003c/em\u003e were programmed to deliver messages to users, these systems worked to optimize workflow in trauma centers and continuously serve as coaching AI, respectively [16, 20]. Moving forward, the reinforcement learning-based virtual AI assistant for children with Down syndrome continuously refined its therapeutic goals by interpreting user feedback and optimizing future interactions [21]. In radiotherapy, \u003cem\u003eGPT-Plan\u003c/em\u003e exemplifies goal-directed reasoning by iteratively adjusting treatment plans through self-corrective loops, mimicking the adaptive deliberation of expert dosimetrists [22], and \u003cem\u003eMultiMedRes\u003c/em\u003e extends this paradigm even further with proactive goal decomposition, breaking complex diagnostic challenges into adaptive subgoals that evolve as new information emerges [24]. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDespite these advances, many current systems still display goal-directed and adaptive behavior only within narrow, task-specific boundaries, often constrained by rule-based logic. Expanding this capability to support temporal adaptability, learning, redefining goals across patient trajectories, and shifting clinical evidence remains a critical frontier. Such longitudinal goal management is particularly vital in healthcare, where long-term, personalized care depends on systems that can dynamically reprioritize objectives and strategies as conditions evolve.\u003c/p\u003e\n\u003cp\u003ed) \u0026nbsp; \u003cem\u003eAction Initiation\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eAnother hallmark of agentic AI is the ability to independently trigger and coordinate actions, transforming them from passive decision aids into active participants in clinical workflows. Modern systems increasingly demonstrate this by autonomously invoking external tools, integrating resources, and orchestrating multi-agent collaboration. \u003cem\u003eChatExosome\u003c/em\u003e, for example, employs retrieval-augmented generation to autonomously query specialized databases and literature, generating evidence-based diagnostic insights without human prompting [25]. Similarly, \u003cem\u003eProtChat\u003c/em\u003e integrates GPT-4 with domain-specific protein language models, generating JSON outputs and seamlessly coordinating multiple analytic tools within a multi-agent framework to perform complex protein property predictions and interaction analyses [23]. The \u003cem\u003eMultiMedRes\u003c/em\u003e framework demonstrates even more advanced orchestration, with its learner agent dynamically querying and integrating domain expert models to iteratively refine multimodal medical reasoning [24]. \u003cem\u003eGPT-Plan\u003c/em\u003e demonstrates multi-agent collaboration in clinical workflows by autonomously initiating optimization processes and leveraging historical treatment data to enhance radiotherapy plans [22].\u003c/p\u003e\n\u003cp\u003eThese capabilities illustrate a significant shift from passive decision support systems to active, self-directed agents that can mobilize knowledge and computational resources in real-time. In contrast, action-initiation mechanisms are absent in some highly autonomous diagnostic systems for conditions like diabetic retinopathy [41-44] or colon polyps [45], where generating a diagnosis or referral typically does not require invoking external tools or APIs. Nevertheless, this growing capacity for action-initiating autonomy raises essential concerns regarding control boundaries, auditability, and ensuring that all autonomous actions remain aligned with clinical objectives and safety standards. As these systems gain greater procedural independence, careful governance will be essential to mitigate risks and ensure their safe integration into healthcare environments.\u003c/p\u003e\n\u003cp\u003ee) \u0026nbsp; \u003cem\u003eCaution and constraints\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eBesides all the current applications mentioned, the promising future era in which Agentic AI is expected to contribute significantly to healthcare [34], it is not a given that they will deliver better results. \u0026nbsp;For example, although \u003cem\u003eSmartText\u003c/em\u003e exhibited better workflow automation, the results were inferior to those of \u003cem\u003eMyCoach\u003c/em\u003e, the other intervention in the study [16]. This system was not an agentic AI because it relied on participants\u0026rsquo; spoken intent to deliver coaching responses. This may suggest that providing patients with medical advice is more effective if they actively seek it, at least for lifestyle changes.\u003c/p\u003e\n\u003cp\u003eMoreover, additional warnings in other fields need to be highlighted, raising concerns about the future of healthcare applications. Recent literature provides substantial evidence that agentic AI, despite its promise, has significant limitations in specific contexts. For tasks with predictable, fixed workflows, agentic AI\u0026apos;s computational cost and latency make it unsuitable for real-time applications, while simpler automation frameworks offer greater efficiency [31]. In scenarios with low error tolerance, the susceptibility of agentic AI to hallucinations and incorrect reasoning poses unacceptable risks in critical decision-making contexts [38]. The \u0026quot;black box\u0026quot; nature of these systems, referred to as the non-explainability of their reasoning or internal workflow process [46], presents significant challenges. This lack of transparency can pose compliance risks and erode stakeholder trust in regulated industries, such as healthcare and finance [47]. Cybersecurity concerns are equally pressing because of vulnerabilities to manipulation, prompt injection attacks, and data privacy risks [38, 47].\u003c/p\u003e\n\u003cp\u003eAdditionally, when predictable outcomes are essential, autonomous decision-making may be problematic for agentic AI applications, such as legal contracts, where precision and predictability are paramount [48]. These constraints collectively suggest that agentic AI should be deployed selectively, carefully considering task requirements and risk tolerance. Using agentic AI systems in routine, predictable workflows would be like using a sledgehammer to crack a nut; hence, they are best suited for complex, dynamic tasks that demand adaptability and autonomous decision-making.\u003c/p\u003e\n\u003cp\u003ef) \u0026nbsp; \u0026nbsp;\u003cem\u003eEthical considerations\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eAgentic AI is widely praised for its autonomy and potential to support clinicians in ethically complex decisions. Still, this very autonomy can also increase the risk of overreliance on systems that may not always be fully understood or properly regulated. A recent example is the Bioethics Artificial Intelligence Advisory (BAIA) framework introduced by Roy et al. (2025), which aims to assist with moral decision-making in healthcare. However, in such high-stakes situations, agentic systems must be more than technically capable, they must be built on high-quality, diverse data, routinely audited for bias and fairness, and designed to explain their reasoning in ways that healthcare professionals and families can clearly understand [49], avoiding to adopt a \u0026ldquo;black box\u0026rdquo; model at all costs [46]. Without these safeguards, critical questions will inevitably arise: Who is responsible if an autonomous AI causes harm? How should clinicians share authority with a system that acts independently? AI must be grounded in transparency and evidence-based design to address these concerns, ensuring accountability, patient safety, and trust in high-stakes clinical settings [50, 51].\u003c/p\u003e\n\u003cp\u003eg) \u0026nbsp; \u003cem\u003eStrengths and limitations\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eThis review identifies several key limitations in the included studies. Most were narrowly focused, used structured or simulated data, and lacked real-world clinical validation, with only one study involving actual patients [16]. Although highly agentic, the systems failed to demonstrate the ultimate level of agency according to current literature [3, 6, 11], presenting with no adaptive behavior over time and lacking long-term memory for other clinical settings. The few studies reflect the limited application of agentic AI in healthcare, and publication bias may have further restricted the evidence base. Additionally, inconsistent terminology across studies underscores the need for more precise definitions and standardized evaluation criteria [5]. However, this systematic review also provides essential strengths. It is the first comprehensive evaluation of agentic AI systems in healthcare, offering clear conceptual definitions, identifying their operational features, and synthesizing the current state of both single- and multi-agent architectures across diverse clinical domains. Besides, it aligns the analysis with practical constraints that must be addressed for safe and effective clinical adoption.\u003c/p\u003e\n\u003cp\u003eh) \u0026nbsp; \u003cem\u003eFuture Directions\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eA multidisciplinary approach is essential to enhance system robustness, clinical relevance, and human-AI collaboration. Targeted use in high-demand areas, such as screening or chronic care, may offer the most immediate value. Still, broader integration will require infrastructure support, legal clarity, and ongoing research in this rapidly evolving field.\u003c/p\u003e\n\u003cp\u003eThe development of agentic AI systems in healthcare is still in its early stages, but the results of this review suggest a clear path toward more mature and impactful applications. Therefore, future work should focus on advancing these systems from proof-of-concept and simulation into real-world clinical environments, guided by structured benchmarks, multidisciplinary evaluation, and robust validation frameworks necessary to address all the gaps related to implementing AI agents.\u003c/p\u003e\n\u003cp\u003eHence, future development of agentic AI should focus on moving beyond experimental systems toward real-world, clinically integrated tools that function as intelligent collaborators rather than passive assistants [51]. These systems must incorporate transparency mechanisms and strengthen human-AI collaboration, ensuring they support rather than replace clinicians [52]. Additionally, the creation of holistic benchmarks is essential to evaluate technical performance, explainability, usability, and trust from a human-centered perspective, ensuring safety and real-world applicability [53]. Further randomized controlled trials with more sophisticated agent-based systems, extended follow-up periods, and larger, more diverse patient populations are necessary to support these advancements.\u003c/p\u003e"},{"header":"5. CONCLUSION","content":"\u003cp\u003eIn conclusion, although the development of agentic AI systems in healthcare is still in its early stages, they represent a transformative step in the evolution of healthcare technologies. These systems offer the potential to autonomously support clinical decision-making, optimize workflows, and enhance patient care through adaptive, goal-driven behaviors. Early applications show promise across diagnostics, treatment planning, and patient engagement; however, their real-world use remains limited and largely experimental, and standardized definitions are still needed. Ongoing research must focus on rigorous validation, ethical design, human-AI collaboration, and seamless integration into clinical environments to realize its full impact. As the field advances, agentic AI holds the potential to become a valuable partner in medicine, one that not only performs tasks but understands, learns, and acts in service of better health outcomes. To make this vision a reality, future efforts must also address practical challenges such as computational cost and latency, which can limit their suitability in time-sensitive clinical contexts. Enhancing processing efficiency, simplifying agent coordination, and leveraging technologies like edge computing will ensure these systems are effective and scalable in real-world healthcare settings.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eACKNOWLEDGEMENTS: The figures were created in BioRender. Collaco, B. (2025) https://BioRender.com/tkl5pm7. This study was funded by Mayo Clinic and the generosity of Eric and Wendy Schmidt; Dalio Philanthropies; and Gerstner Philanthropies. The funder played no role in study design, data collection, analysis and interpretation of data, or the writing of this manuscript.\u003c/p\u003e\n\u003cp\u003eAll authors declare no financial or non-financial competing interests.\u003c/p\u003e\n\u003cp\u003eAUTHOR CONTRIBUTION: BC and AF helped with the conceptualization, and BC, AG, SH contributed to the methodology, including study screening, selection, and data extraction. Validation and formal analysis of the results was performed by BC. The original draft was prepared by BC and SH, with review and editing by SP, CG, AG, and AF. Supervision and project administration was the responsibility of AF. All authors read and approved of the final manuscript.\u003c/p\u003e\n\u003cp\u003eDATA AVAILABILITY: The data supporting the findings of this systematic review are available within the article.\u003c/p\u003e\n\u003cp\u003eCODE AVAILABILITY: Not applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003ePressman, S.M., et al., \u003cem\u003eAI and Ethics: A Systematic Review of the Ethical Considerations of Large Language Model Use in Surgery Research\u003c/em\u003e. Healthcare (Basel), 2024. 12(8).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWilcox, A., M. Griffith, and O. Griffith, \u003cem\u003eLooking Forward to AI and Medicine: Where Are We, and Where Are We Going?\u003c/em\u003e Mo Med, 2025. 122(1): p. 34\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKarunanayake, N., \u003cem\u003eNext-generation agentic AI for transforming healthcare\u003c/em\u003e. Informatics and Health, 2025. 2(2): p. 73\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWooldridge, M. and N.R. Jennings, \u003cem\u003eIntelligent agents: theory and practice\u003c/em\u003e. The Knowledge Engineering Review, 1995. 10(2): p. 115\u0026ndash;152.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHughes, L., et al., \u003cem\u003eAI Agents and Agentic Systems: A Multi-Expert Analysis.\u003c/em\u003e Journal of Computer Information Systems: p. 1\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAcharya, D.B., K. Kuppan, and B. Divya, \u003cem\u003eAgentic AI: Autonomous Intelligence for Complex Goals - A Comprehensive Survey\u003c/em\u003e. IEEE Access, 2025. 13: p. 18912\u0026ndash;18936.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRuan, J., et al., \u003cem\u003eTPTU: large language model-based AI agents for task planning and tool usage.\u003c/em\u003e arXiv preprint arXiv:2308.03427, 2023.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMaleki Varnosfaderani, S. and M. Forouzanfar, \u003cem\u003eThe Role of AI in Hospitals and Clinics: Transforming Healthcare in the 21st Century\u003c/em\u003e. Bioengineering (Basel), 2024. 11(4).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePage, M.J., et al., \u003cem\u003eThe PRISMA 2020 statement: an updated guideline for reporting systematic reviews\u003c/em\u003e. BMJ, 2021. 372: p. n71.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCollaco, B. \u003cem\u003eTHE ROLE OF AGENTIC AI IN HEALTHCARE: A SYSTEMATIC REVIEW. PROSPERO CRD420251043917\u003c/em\u003e. 2025; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.crd.york.ac.uk/PROSPERO/view/CRD420251043917\u003c/span\u003e\u003cspan address=\"https://www.crd.york.ac.uk/PROSPERO/view/CRD420251043917\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShavit, Y., et al., \u003cem\u003ePractices for governing agentic AI systems\u003c/em\u003e. Research Paper, OpenAI, 2023.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJheng, Y.C., et al., \u003cem\u003eThe era of artificial intelligence-based individualized telemedicine is coming\u003c/em\u003e. J Chin Med Assoc, 2020. 83(11): p. 981\u0026ndash;983.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGenovese, A., et al., \u003cem\u003eArtificial intelligence in clinical settings: a systematic review of its role in language translation and interpretation\u003c/em\u003e. Ann Transl Med, 2024. 12(6): p. 117.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLight, M., \u003cem\u003eEndNote 1-2-3 Easy! Reference Management for the Professional, Second Edition\u003c/em\u003e, A. Agrawal, 2009, \u003cem\u003eSpringer, New York, USA, Price: \u0026euro;44.95, Soft Cover, 294 pages, ISBN: 978-0-387-95900-9, Website: www.springer.com.\u003c/em\u003e South African Journal of Botany, 2010. 76.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSterne, J.A., et al., \u003cem\u003eROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions\u003c/em\u003e. BMJ, 2016. 355: p. i4919.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHassoon, A., et al., \u003cem\u003eRandomized trial of two artificial intelligence coaching interventions to increase physical activity in cancer survivors\u003c/em\u003e. NPJ digital medicine, 2021. 4(1): p. 168.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHiggins, J.P., et al., \u003cem\u003eThe Cochrane Collaboration's tool for assessing risk of bias in randomised trials\u003c/em\u003e. BMJ, 2011. 343: p. d5928.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCardoso, R., et al., \u003cem\u003eAn updated meta-analysis of novel oral anticoagulants versus vitamin K antagonists for uninterrupted anticoagulation in atrial fibrillation catheter ablation\u003c/em\u003e. Heart Rhythm, 2018. 15(1): p. 107\u0026ndash;115.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcGuinness, L.A. and J.P.T. Higgins, \u003cem\u003eRisk-of-bias VISualization (robvis): An R package and Shiny web app for visualizing risk-of-bias assessments\u003c/em\u003e. Research Synthesis Methods, 2021. 12(1): p. 55\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCroatti, A., et al., \u003cem\u003eBDI personal medical assistant agents: The case of trauma tracking and alerting.\u003c/em\u003e Artificial Intelligence in Medicine, 2019. 96((Croatti A.,
[email protected]; Montagna S.,
[email protected]; Ricci A.,
[email protected]) Computer Science and Engineering Department (DISI), University of Bologna, Campus of Cesena, Via dell'Universit\u0026agrave; 50, Cesena, Italy): p. 187\u0026ndash;197.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMariselvam, J., S. Rajendran, and Y. Alotaibi, \u003cem\u003eReinforcement learning-based AI assistant and VR play therapy game for children with Down syndrome bound to wheelchairs\u003c/em\u003e. AIMS Mathematics, 2023. 8(7): p. 16989\u0026ndash;17011.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang, Q., et al., \u003cem\u003eA feasibility study of automating radiotherapy planning with large language model agents\u003c/em\u003e. Physics in medicine and biology, 2025. 70(7).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHuang, H., et al., \u003cem\u003eProtChat: An AI Multi-Agent for Automated Protein Analysis Leveraging GPT-4 and Protein Language Model\u003c/em\u003e. Journal of chemical information and modeling, 2025. 65(1): p. 62\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGu, Z., et al., \u003cem\u003eA Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning\u003c/em\u003e. Advanced Intelligent Systems, 2025.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang, Z., et al., \u003cem\u003eChatExosome: An Artificial Intelligence (AI) Agent Based on Deep Learning of Exosomes Spectroscopy for Hepatocellular Carcinoma (HCC) Diagnosis.\u003c/em\u003e Analytical chemistry, 2025. 97(8): p. 4643\u0026ndash;4652.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHassoon, A., et al., \u003cem\u003eIncreasing Physical Activity Amongst Overweight and Obese Cancer Survivors Using an Alexa-Based Intelligent Agent for Patient Coaching: Protocol for the Physical Activity by Technology Help (PATH) Trial\u003c/em\u003e. JMIR Res Protoc, 2018. 7(2): p. e27.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOgbu, D., \u003cem\u003eAgentic AI in Computer Vision Domain - Recent Advances and Prospects\u003c/em\u003e. International Journal of Research Publication and Reviews, 2024. 5(11): p. 5102\u0026ndash;5120.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLee, H., \u003cem\u003eThe rise of ChatGPT: Exploring its potential in medical education\u003c/em\u003e. Anat Sci Educ, 2024. 17(5): p. 926\u0026ndash;931.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHadi, M.U., et al., \u003cem\u003eLarge language models: a comprehensive survey of its applications, challenges, limitations, and future prospects\u003c/em\u003e. Authorea Preprints, 2023. 1: p. 1\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSanderson, K., \u003cem\u003eGPT-4 is here: what scientists think\u003c/em\u003e. Nature, 2023. 615(7954): p. 773.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSapkota, R., K.I. Roumeliotis, and M. Karkee, \u003cem\u003eAI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenge.\u003c/em\u003e arXiv preprint arXiv:2505.10468, 2025.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSchneider, J., \u003cem\u003eGenerative to Agentic AI: Survey, Conceptualization, and Challenges.\u003c/em\u003e arXiv preprint arXiv:2504.18875, 2025.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCastelfranchi, C., \u003cem\u003eModelling social action for AI agents\u003c/em\u003e. Artificial Intelligence, 1998. 103(1): p. 157\u0026ndash;182.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHosseini, S. and H. Seilani, \u003cem\u003eThe role of agentic AI in shaping a smart future: A systematic review\u003c/em\u003e. Array, 2025. 26.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFischer, K., et al. \u003cem\u003eSophisticated and distributed: The transportation domain\u003c/em\u003e. in \u003cem\u003eProceedings of 9th IEEE Conference on Artificial Intelligence for Applications\u003c/em\u003e. 1993.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu, G., et al., \u003cem\u003eWireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception.\u003c/em\u003e arXiv preprint arXiv:2505.23275, 2025.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eViswanathan, S.P.a.M. \u003cem\u003e5 Levels of agentic AI intelligence for enterprise use\u003c/em\u003e. Outshift 2025 05/26/2025]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://outshift.cisco.com/blog/agentic-ai-intelligence-for-enterprise-use\u003c/span\u003e\u003cspan address=\"https://outshift.cisco.com/blog/agentic-ai-intelligence-for-enterprise-use\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e\u003cem\u003eThe Six Levels of Agentic Behavior\u003c/em\u003e. 2025 05/24/25]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.vellum.ai/blog/levels-of-agentic-behavior\u003c/span\u003e\u003cspan address=\"https://www.vellum.ai/blog/levels-of-agentic-behavior\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDemazeau, Y. and J.P. M\u0026uuml;ller, \u003cem\u003eDecentralized A.I\u003c/em\u003e. 1990: North Holland.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChan, A., et al., \u003cem\u003eHarms from Increasingly Agentic Algorithmic Systems\u003c/em\u003e, in \u003cem\u003e2023 ACM Conference on Fairness Accountability and Transparency\u003c/em\u003e. 2023. p. 651\u0026ndash;666.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbr\u0026agrave;moff, M.D., et al., \u003cem\u003eMitigation of AI adoption bias through an improved autonomous AI system for diabetic retinal disease\u003c/em\u003e. NPJ digital medicine, 2024. 7(1): p. 369.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWolf, R.M., et al., \u003cem\u003eAutonomous artificial intelligence increases screening and follow-up for diabetic retinopathy in youth: the ACCESS randomized control trial\u003c/em\u003e. Nature communications, 2024. 15(1): p. 421.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbr\u0026agrave;moff, M.D., et al., \u003cem\u003ePivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices\u003c/em\u003e. NPJ digital medicine, 2018. 1: p. 39.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbramoff, M.D., et al., \u003cem\u003eAutonomous artificial intelligence increases real-world specialist clinic productivity in a cluster-randomized trial\u003c/em\u003e. NPJ digital medicine, 2023. 6(1): p. 184.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDjinbachian, R., et al., \u003cem\u003eAUTONOMOUS ARTIFICIAL INTELLIGENCE VERSUS AI ASSISTED HUMAN OPTICAL DIAGNOSIS OF COLORECTAL POLYPS: a RANDOMIZED CONTROLLED TRIAL\u003c/em\u003e. Gastrointestinal endoscopy, 2024. 99(6): p. AB11.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMarcus, E. and J. Teuwen, \u003cem\u003eArtificial intelligence and explanation: How, why, and when to explain black boxes\u003c/em\u003e. Eur J Radiol, 2024. 173: p. 111393.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSimoes, B.S. \u003cem\u003eAgentic AI: A Promising Evolution, But Not Without Limits\u003c/em\u003e. 2025 05/26/25]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.blueprintsys.com/blog/agentic-ai-a-promising-evolution-but-not-without-limits\u003c/span\u003e\u003cspan address=\"https://www.blueprintsys.com/blog/agentic-ai-a-promising-evolution-but-not-without-limits\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWarmly. \u003cem\u003eAI Agentic Workflows: Definitions, Use Cases \u0026amp; Software\u003c/em\u003e. 2025 05/26/25]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.warmly.ai/p/blog/ai-agentic-workflows\u003c/span\u003e\u003cspan address=\"https://www.warmly.ai/p/blog/ai-agentic-workflows\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDutta Roy, T.P., \u003cem\u003eBioethics Artificial Intelligence Advisory (BAIA): An Agentic Artificial Intelligence (AI) Framework for Bioethical Clinical Decision Support\u003c/em\u003e. Cureus, 2025. 17(3): p. e80494.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePapagni, G., et al., \u003cem\u003eArtificial agents\u0026rsquo; explainability to support trust: considerations on timing and context\u003c/em\u003e. Ai \u0026amp; Society, 2022. 38(2): p. 947\u0026ndash;960.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePapagni, G. and S. Koeszegi, \u003cem\u003eA pragmatic approach to the intentional stance semantic, empirical and ethical considerations for the design of artificial agents\u003c/em\u003e. Minds and Machines, 2021. 31(4): p. 505\u0026ndash;534.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHeer, J., \u003cem\u003eAgency plus automation: Designing artificial intelligence into interactive systems.\u003c/em\u003e Proceedings of the National Academy of Sciences, 2019. 116(6): p. 1844\u0026ndash;1850.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGridach, M., et al., \u003cem\u003eAgentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions.\u003c/em\u003e arXiv preprint arXiv:2503.08979, 2025.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1. Baseline characteristics of included studies\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" align=\"\" width=\"124%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eStudy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eCountry\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eStudy Type\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eDomain\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"5\" valign=\"top\" style=\"width: 70px;\"\u003e\n \u003cp\u003eAgentic AI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eName\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eMain Clinical Task\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eKey Findings\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003eFeatures\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003eLimitations\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eCroatti 2019 [20]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eItaly\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eEmergency Medicine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eTraumaTracker\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eReal-time trauma tracking and generation of clinical alerts during resuscitation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eImproved documentation quality, enabled real-time alerts aligned with trauma team workflows\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomous trauma event tracking;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior for optimizing workflow with multimodal interaction: tablet, smart glasses;\u003c/p\u003e\n \u003cp\u003e- Real-time reports and alert generation\u003c/p\u003e\n \u003cp\u003e- Short-term memory during each trauma case;\u003c/p\u003e\n \u003cp\u003e- Multi-Agent collaboration: two cooperating BDI agents;\u003c/p\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Narrow task scope: trauma only;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on structured resuscitation scenarios;\u003c/p\u003e\n \u003cp\u003e- Tool and integration constraints;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation and generalizability;\u003c/p\u003e\n \u003cp\u003e- Requires manual input for some tasks;\u003c/p\u003e\n \u003cp\u003e- Over-alerting risk in specific rules;\u003c/p\u003e\n \u003cp\u003e- Customization and interoperability improvements needed;\u003c/p\u003e\n \u003cp\u003e- Not adaptive in a dynamic or machine learning sense;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eHassoon 2021\u0026nbsp;[16]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eRCT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eLifestyle Intervention\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eSmartText\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003ePersonalized physical activity coaching to increase daily step counts\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eHad a modest, non-significant effect\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomous message generation;\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior as a coaching system;\u003c/p\u003e\n \u003cp\u003e- Sends three personalized text messages per day;\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior based on participant data;\u003c/p\u003e\n \u003cp\u003e- Short-term memory by storing recent physical activity;\u003c/p\u003e\n \u003cp\u003e- Single Agent system;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Narrow scope: increase footsteps;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Short study duration: 4 weeks;\u003c/p\u003e\n \u003cp\u003e- Small sample size: 42 patients;\u003c/p\u003e\n \u003cp\u003e- Limited generalizability\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e- Not adapted to complex clinical environments;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: structured, rule-driven, and preprogrammed setting, not adaptable over time;\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eMariselvam 2023 [21]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eSaudi Arabia\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003ePediatric rehabilitation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eReinforcement Learning-Based Virtual AI Assistant\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eEnhancing motor, cognitive, and social skills in children with Down syndrome using VR play therapy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003ePPO + Actor-Critic performed best with 87% accuracy in adapting difficulty and supporting skill development; suitable for independent gameplay\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomous decision-making in a simulated setting;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior for skills improvement;\u003c/p\u003e\n \u003cp\u003e- Action initiation by tilting the virtual board, moving the ball, and introducing obstacles;\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior by adjusting game difficulty, analyzing performance, and modifying strategies;\u003c/p\u003e\n \u003cp\u003e- Short-term memory during training and gameplay;\u003c/p\u003e\n \u003cp\u003e- Single Agent system;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e\u0026nbsp;- Narrow scope: single VR therapy game;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on sub-modules;\u003c/p\u003e\n \u003cp\u003e- Lack of physical agency;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation and generalizability: children who used wheelchairs with Down syndrome;\u003c/p\u003e\n \u003cp\u003e- No real-world autonomy;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: restricted to a single game environment, not adaptable over time;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eGu 2025\u0026nbsp;[24]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eRadiology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eMultiMedRes\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eDVQA between sequential chest X-rays\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eAchieved state-of-the-art performance on DVQA tasks; GPT-4-based learner agent outperformed both fine-tuned and fully supervised models; effective in zero-shot settings\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomously determines when to stop querying;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior to address medical multimodal reasoning problems;\u003c/p\u003e\n \u003cp\u003e- Initiation of external actions by launching API calls on its own;\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior by changing its reasoning strategy;\u003c/p\u003e\n \u003cp\u003e- Short-term, task-specific memory;\u003c/p\u003e\n \u003cp\u003e- Single LLM agent + expert models (APIs);\u003c/p\u003e\n \u003cp\u003e- Retrieval-Augmented Reasoning;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Narrow Scope: DVQA for chest X-rays;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on expert modules;\u003c/p\u003e\n \u003cp\u003e- Lack of physical agency;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation and generalizability;\u003c/p\u003e\n \u003cp\u003e- No real-world autonomy;\u003c/p\u003e\n \u003cp\u003e- Absence of patient-specific context;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: task-specific and within constraints, not adaptable over time;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eHuang 2025 [23]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eChina\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eBiomedicine\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eProtChat\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eAutomation of complex protein analysis tasks using GPT-4 and PLLMs\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eAccurately automated protein analysis tasks with strong performance across benchmark datasets, using standard evaluation metrics, such as Pearson, RMSE and MAE, ROC curve, and PR curve.\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e\u0026nbsp;- Autonomous task planning and execution via GPT-4;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior for protein understanding tasks;\u003c/p\u003e\n \u003cp\u003e- Initiation of external actions: invokes PLLMs, runs functions, writes JSON results, and passes them to downstream agents\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior by handling multiple tasks;\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e- Short-term, session-bound memory;\u003c/p\u003e\n \u003cp\u003e- Multi-agent collaboration: User Proxy, Inference, Evaluation, Visualization, and Chat Manager agents;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Narrow scope: protein analysis tasks;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on sub-modules;\u003c/p\u003e\n \u003cp\u003e- Lack of physical agency;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation: Only tested on computational protein datasets;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: task-specific and instruction-driven, not adaptable over time;\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eWang 2025\u0026nbsp;[22]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eChina\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eRadiation Oncology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eGPT-Plan\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eOptimization and generation of radiotherapy treatment plans\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eMatched or outperformed expert human planners and auto-planning systems in plan quality and efficiency, especially effective in OAR sparing and optimization iterations for lung and cervical cancer\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomous plan generation;\u003c/p\u003e\n \u003cp\u003e- Goal-directed behavior for treatment optimization;\u003c/p\u003e\n \u003cp\u003e- Initiates external actions/tool calls\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior by refining treatment plans;\u003c/p\u003e\n \u003cp\u003e- Multi-agent collaboration: Dosimetrist, Physicist, TPS_proxy, Human_proxy;\u003c/p\u003e\n \u003cp\u003e- Retrieval-Augmented Reasoning: Retriever tool;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Narrow scope: lung and cervical cancer;\u003c/p\u003e\n \u003cp\u003e- Limited domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on sub-modules;\u003c/p\u003e\n \u003cp\u003e- Lack of physical agency;\u003c/p\u003e\n \u003cp\u003e- Small sample size: 17 patients;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation and generalizability;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: not across sessions, not adaptable over time;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eYang 2025 [25]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 5px;\"\u003e\n \u003cp\u003eChina\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eExperimental\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 8px;\"\u003e\n \u003cp\u003eOncology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 7px;\"\u003e\n \u003cp\u003eChatExosome\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003eDiagnose hepatocellular carcinoma from plasma exosome Raman spectroscopy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003eChatExosome achieved 94.1% accuracy in distinguishing HCC from controls, including 87.5% in AFP-negative HCC cases\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18px;\"\u003e\n \u003cp\u003e- Autonomous execution of multi-step reasoning;\u003c/p\u003e\n \u003cp\u003e- Goal-directed loop and decision-making;\u003c/p\u003e\n \u003cp\u003e- External tool invocation (FFT classifier, RAG retrieval) in response to file or text inputs;\u003c/p\u003e\n \u003cp\u003e- Adaptive behavior by switching between predefined models\u003c/p\u003e\n \u003cp\u003e- Short-term memory from a predefined external knowledge base;\u003c/p\u003e\n \u003cp\u003e- Single Agent system\u003c/p\u003e\n \u003cp\u003e-\u0026nbsp;Retrieval-Augmented Reasoning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e- Partially narrow scope: demonstrated capacity of more than one type of cancer diagnosis;\u003c/p\u003e\n \u003cp\u003e- Single domain;\u003c/p\u003e\n \u003cp\u003e- Reliance on pretrained model;\u003c/p\u003e\n \u003cp\u003e- Lack of physical agency;\u003c/p\u003e\n \u003cp\u003e- Limited clinical validation and generalizability;\u003c/p\u003e\n \u003cp\u003e- Limited adaptability: does not self-update; not adaptable over time;\u003c/p\u003e\n \u003cp\u003e- No long-term memory;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Agentic AI, agentic systems, AI agents, autonomous systems, healthcare","lastPublishedDoi":"10.21203/rs.3.rs-7234499/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7234499/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground/ Objectives:\u003c/h2\u003e\u003cp\u003eAgentic AI represents a promising evolution of AI technology applied to healthcare, with systems increasingly capable of operating autonomously to achieve defined clinical goals. However, the literature lacks conceptual clarity between \u0026ldquo;AI agents\u0026rdquo; and \u0026ldquo;agentic AI\u0026rdquo;, and few studies have rigorously explored their clinical applications. Therefore, this study aims to conduct a novel systematic review addressing this gap by examining agentic AI systems in healthcare settings, characterizing their applications, features, outcomes, and limitations, and clarifying the conceptual distinctions between AI agents and agentic AI systems using predefined and objective criteria.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e\u003cp\u003eA comprehensive search was conducted across PubMed, Embase, Cochrane, Scopus, and Google Scholar on April 6th, 2025. Studies were included if they involved AI systems in healthcare settings that demonstrated the following agentic features: autonomous operation, goal-directed behavior, and initiating action. Data on the clinical tasks achieved by the agents, key findings, features, and limitations were collected from the included studies. Screening and extraction followed PRISMA guidelines, with Risk of Bias assessed using ROBINS-I and Cochrane's Risk of Bias tools.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003eOf 984 retrieved records, seven studies met the inclusion criteria, spanning domains such as emergency medicine, oncology, radiology, and rehabilitation. Multi-agent architecture was frequently used to decompose and coordinate complex workflows. Among the included studies, the AIs showed high accuracy in diagnosing cancer patients, conducting treatment plans, sending alerts, coaching messages, analysing image data, and adapting to challenging experimental scenarios. While demonstrating potential for improved efficiency, task accuracy, and patient engagement, significant limitations were noted: narrow task scope, lack of physical agency, limited clinical validation, and barriers to integration into real-world healthcare systems. Only one system had been deployed in a patient-facing trial setting.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e\u003cp\u003eThe current literature suggests an emerging role and application of Agentic AI, holding promise with the potential to revolutionize diagnostics, triage, treatment planning, and patient management. However, real-world implementation and evaluations in the literature are limited. Future research must address critical validation, regulation, ethics, and clinical integration challenges to realize their full potential. Clear operational definitions and frameworks for evaluating agency are essential to support safe and effective deployment of these systems.\u003c/p\u003e","manuscriptTitle":"The Role of Agentic Artificial Intelligence in Healthcare: A Systematic Review","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-07 07:09:57","doi":"10.21203/rs.3.rs-7234499/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-09-13T11:22:15+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-12T14:09:19+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-11T00:10:27+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-13T09:47:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"174801134484670304724036032686179352765","date":"2025-08-09T03:26:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"83323715927987122940114747589948171197","date":"2025-08-08T19:03:18+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"196388882340901148888602275157730467339","date":"2025-08-06T08:30:54+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"103018505470055545394882924326888412086","date":"2025-08-06T01:53:13+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-05T01:51:24+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-08-05T00:57:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-08-04T08:41:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-07-28T13:31:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b0d87884-1565-49c3-999e-a53666ffa4f0","owner":[],"postedDate":"August 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":52739242,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":52739243,"name":"Health sciences/Health care"},{"id":52739244,"name":"Physical sciences/Mathematics and computing"},{"id":52739245,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-03-16T16:03:12+00:00","versionOfRecord":{"articleIdentity":"rs-7234499","link":"https://doi.org/10.1038/s41746-026-02517-5","journal":{"identity":"npj-digital-medicine","isVorOnly":false,"title":"npj Digital Medicine"},"publishedOn":"2026-03-14 15:59:39","publishedOnDateReadable":"March 14th, 2026"},"versionCreatedAt":"2025-08-07 07:09:57","video":"","vorDoi":"10.1038/s41746-026-02517-5","vorDoiUrl":"https://doi.org/10.1038/s41746-026-02517-5","workflowStages":[]},"version":"v1","identity":"rs-7234499","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7234499","identity":"rs-7234499","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.