Evaluating the Clinical Competence of Artificial Intelligence Applications in Psychiatry: A Systematic Review

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Artificial intelligence (AI) is an increasingly promising technology in psychiatry with the potential to transform mental healthcare. The number of AI applications designed for psychiatric diagnosis and treatment has grown substantially over the last few years; however, clinicians remain concerned about the real-world readiness of these applications in clinical care settings. These concerns have some validity, given recently reported cases of worsening psychosis and suicide attempts associated with AI application use. While there are a few studies that have examined the performance of AI applications on various multiple-choice question banks, the clinical usefulness and practical application of such test performance have yet to be assessed. Assessing test performance will provide an appraisal of the case complexity and appropriateness of the clinical decision-making process used by AI applications. Such an assessment offers a more definitive understanding of the real-world readiness of various AI applications in clinical scenarios, as well as their comparative performance relative to that of human physicians. We conducted a systematic review of peer-reviewed publications from April 2014 to July 2025, evaluating the real-world clinical readiness of publicly available psychiatric AI applications. Following PRISMA and MOOSE guidelines, 68 publications were identified, of which 24 met the inclusion criteria, yielding 24 unique applications. Each AI application was evaluated based on the test administered for clinical competence, with case complexity rated using the Amsterdam Clinical Challenge Scale and clinical decision-making assessed via Miller’s Pyramid of Clinical Competence. The inclusion of the various domains of psychiatric practice—evaluation, diagnosis, psychopharmacology, psychotherapy, and psychosocial intervention was also taken into consideration. Competence was determined based on the intersection of case complexity and decision-making level.Findings revealed substantial heterogeneity in performance across the various domains of psychiatry, with no AI application achieving human-level competence in all domains of psychiatric practice. Some psychotherapy-focused tools demonstrated moderate complexity and comparatively higher competence, while most applications remained limited to basic knowledge application (Miller Level 2). Core clinical domains, such as evaluation and diagnosis, were underdeveloped, with more than 70% of tools lacking defined complexity or competence ratings. Importantly, none of the applications used board certification-style testing; only 10.5% reported training on board-level material, and just 4.2% disclosed the use of training datasets.These results highlight a critical gap between the potential of psychiatric AI and its demonstrated real-world readiness. Standardized, board-style evaluation, transparency in training data, and more rigorous measures of clinical decision-making are essential to building trust and supporting safe integration into psychiatric care.
Full text 90,066 characters · extracted from preprint-html · click to expand
Evaluating the Clinical Competence of Artificial Intelligence Applications in Psychiatry: A Systematic Review | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Evaluating the Clinical Competence of Artificial Intelligence Applications in Psychiatry: A Systematic Review Payton Colantonio, Katarina Milosavljevic, Camille Mero, Mrinal Sharma, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8206381/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Artificial intelligence (AI) is an increasingly promising technology in psychiatry with the potential to transform mental healthcare. The number of AI applications designed for psychiatric diagnosis and treatment has grown substantially over the last few years; however, clinicians remain concerned about the real-world readiness of these applications in clinical care settings. These concerns have some validity, given recently reported cases of worsening psychosis and suicide attempts associated with AI application use. While there are a few studies that have examined the performance of AI applications on various multiple-choice question banks, the clinical usefulness and practical application of such test performance have yet to be assessed. Assessing test performance will provide an appraisal of the case complexity and appropriateness of the clinical decision-making process used by AI applications. Such an assessment offers a more definitive understanding of the real-world readiness of various AI applications in clinical scenarios, as well as their comparative performance relative to that of human physicians. We conducted a systematic review of peer-reviewed publications from April 2014 to July 2025, evaluating the real-world clinical readiness of publicly available psychiatric AI applications. Following PRISMA and MOOSE guidelines, 68 publications were identified, of which 24 met the inclusion criteria, yielding 24 unique applications. Each AI application was evaluated based on the test administered for clinical competence, with case complexity rated using the Amsterdam Clinical Challenge Scale and clinical decision-making assessed via Miller’s Pyramid of Clinical Competence. The inclusion of the various domains of psychiatric practice—evaluation, diagnosis, psychopharmacology, psychotherapy, and psychosocial intervention was also taken into consideration. Competence was determined based on the intersection of case complexity and decision-making level. Findings revealed substantial heterogeneity in performance across the various domains of psychiatry, with no AI application achieving human-level competence in all domains of psychiatric practice. Some psychotherapy-focused tools demonstrated moderate complexity and comparatively higher competence, while most applications remained limited to basic knowledge application (Miller Level 2). Core clinical domains, such as evaluation and diagnosis, were underdeveloped, with more than 70% of tools lacking defined complexity or competence ratings. Importantly, none of the applications used board certification-style testing; only 10.5% reported training on board-level material, and just 4.2% disclosed the use of training datasets. These results highlight a critical gap between the potential of psychiatric AI and its demonstrated real-world readiness. Standardized, board-style evaluation, transparency in training data, and more rigorous measures of clinical decision-making are essential to building trust and supporting safe integration into psychiatric care. Scientific community and society/Business and industry Health sciences/Diseases Health sciences/Health care Health sciences/Medical research Biological sciences/Psychology Social science/Psychology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Introduction Artificial Intelligence (AI) refers to computational systems capable of performing tasks that traditionally require human intelligence, such as pattern recognition, problem-solving, and decision-making. In healthcare, AI applications have rapidly expanded over the last decade due to advances in machine learning algorithms, natural language processing, and increased computational power. AI has demonstrated utility across diverse specialties, including radiology, oncology, dermatology, and cardiology, where it has been applied to diagnostic imaging, treatment planning, and risk stratification, with performance at times comparable to or exceeding that of human clinicians 1,2 . In psychiatry, AI offers a unique promise given the discipline’s reliance on subjective assessments, longitudinal monitoring, and multimodal data integration. Within psychiatric practice, AI encompasses tools such as natural language processing models for clinical documentation and diagnostic support, conversational agents for delivering psychotherapy, predictive algorithms for relapse or suicide risk, and decision support systems for psychopharmacological treatment. Examples include smartphone-based monitoring of mood and behavior, virtual therapeutic chatbots, and large language models (LLMs) capable of synthesizing diagnostic criteria from structured and unstructured data. Despite this potential, the adoption of AI in psychiatry remains limited. Systematic reviews of healthcare AI have identified several barriers, including insufficient data transparency, a lack of standardized evaluation frameworks, and limited clinician involvement in the design and validation processes. Unlike radiology or pathology, psychiatry lacks large standardized datasets with objective biomarkers, relying instead on nuanced clinical interviews, patient self-reports, and complex social contexts that challenge algorithmic modelling. Clinicians have expressed skepticism regarding whether current AI systems can adequately capture the subtleties of psychiatric assessment, particularly in culturally diverse populations and comorbid conditions. Furthermore, patients may be hesitant to accept diagnoses or treatment recommendations from AI systems, particularly when such recommendations are generated through opaque, “black box” processes that lack explainability. These challenges highlight why, compared with other medical specialties, psychiatry has been slower to integrate AI technologies, and raise the question of whether existing tools are sufficiently ready for real-world use to function safely and competently in clinical environments. If successfully integrated, however, AI applications have the potential to transform psychiatric care by enhancing diagnostic accuracy, personalizing treatment, and expanding access to care. For example, AI-driven digital phenotyping has been explored to predict depressive episodes based on smartphone use patterns, while conversational agents delivering cognitive behavioral therapy (CBT) have shown efficacy in reducing symptoms of anxiety and depression 7,8 . Predictive algorithms are also being developed to forecast treatment response trajectories and inform psychopharmacological decisions 9 . These applications suggest that AI could serve as a valuable adjunct to psychiatrists, augmenting, rather than replacing, clinical judgment. The challenges of integrating AI into psychiatric care, despite the promises, are motivated by clinician skepticism. Current AI systems in psychiatry often lack rigorous validation against gold-standard clinical benchmarks, such as board certification-style examinations or structured clinical assessments. The current methods of testing LLM based AI applications for healthcare, for instance, include a variety of question banks drawn from medical student examinations. It is unclear whether AI applications could provide meaningful clinical decisions for real-world clinical scenarios. This lack of real-world, realistic competency testing undermines confidence that AI systems are truly ready for clinical use. Moreover, many tools are developed without meaningful input from psychiatrists during design, limiting their clinical applicability 10 . Issues of transparency in training datasets, heterogeneity in performance across psychiatric domains, and the lack of standardized evaluation protocols remain substantial barriers to trust and implementation 11 . In this study, we present a systematic review of AI applications in psychiatry, with a focus on tests for clinical competence. We rate the tests used by the AI applications for clinical complexity, as well as the sophistication of clinical decisions by these applications. Standardized ratings for clinical complexity (Amsterdam Clinical Challenge Scale) and clinical decision sophistication (Miller’s Pyramid of Clinical Competence) were used. The rating allowed for a more uniform comparison between the various AI applications tested with various question banks. The rating also allowed for comparison against human physicians who are certified using board certification examinations. By systematically analyzing the peer-reviewed literature, this review aims to address clinical skepticism and characterize the current state of AI applications in psychiatry in terms of clinical competence. This is the first known review that seeks to provide a more rigorous understanding of the clinical competence of AI applications in psychiatry, as well as to rank these applications against humans based on their ability to address real-world clinical scenarios. We also seek to highlight gaps in clinical testing and identify areas where rigorous validation is most urgently needed. Understanding these limitations and opportunities is essential to guide future development of AI systems that can achieve true real-world readiness and safely support psychiatric care. Methods A total of 24 AI applications were evaluated for their clinical capabilities in psychiatry across five core domains: psychiatric evaluation, psychiatric diagnosis, psychopharmacology, psychotherapy, and psychosocial interventions. Each application was assessed using Joan’s Briggs Critical Appraisal Tool to ensure transparency, Amsterdam Clinical Challenge Scale (ACCS) to evaluate case complexity, and Miller’s Pyramid of Clinical Competence (MPCC) to determine the depth of clinical reasoning and applied knowledge. We conducted a literature review from March 2024 to July 2025, focusing on peer-reviewed publications that evaluate the clinical capabilities of AI software in Psychiatry. The MOOSE guidelines for Meta-Analysis and Systematic Reviews for Observational Studies were followed to identify AI software accurately. Literature reviews were conducted using PubMed/MEDLINE, Google Scholar, Ovid, JoVE, and ScienceDirect. The research was conducted using search terms including but not limited to “artificial intelligence,” “AI”, “Psychiatry”, “testing”, “board certification”, “examinations”, “Clinical Skills”, “clinical decision”, “clinical capabilities”, “Clinical Competency,” “Performance”, “human comparison”. Search terms also included publicly available Psychiatry AI softwares including but not limited to "ChatGPT DSM V Focus4", "Med-PaLM", "Dragon", "AMIE", "ChatGPT DSM-5 Focused Assistant", "BiolinkBert", "DrugCard", "PubMedGPT", "MindMateGPT", "YourDoctor", "PubMedBert", "Dedicated Psychiatrist", "Psyscribe", "Woebot", "DoctorGPT ALGOMED", "RxGuide (ChatGPT)", "Doctor GPT Chat GPT", "Prognos Application", "Soap note AI", "Ada Health", "Wysa", "Replika", "Talkspace", "Youper", "BetterHelp", "Moodpath aka MindDoc", "Kintsugi", "ReGain", "SuperBetter", "CBT-i Coach", "Happify", "MindShift", "MoodMission", "Headspace", "Calm", "Daylio", "RDMaster", "Limbic AI", "Moodpen", "Calm Mind", "Mind well AI", "Aiberry", "MyClinicalWriter", "Ree AI", "Corti", "Elomia", "Meomind", Sintelly", "DBT Diary Card and Skills Coach", "MoodTools" and "Calm Harm". These search terms were used in all possible combinations. Papers were included if they were published after April 2014, included any of the searched Psychiatry AI software as a subject, and the AI software was publicly available to patients or clinicians. All papers were eligible for inclusion if they included tests on the clinical capabilities of the Psychiatry AI software. Papers involving AI software that were not publicly available for use were excluded from the study. The end users of the AI software (physicians or patients) were not an exclusion criterion in the literature review. The type of AI model used in the AI software and its scope of application were considered, but did not alter the inclusion criteria.* A total of 68 papers were identified through database searching, of which 56 publications remained after duplicates were removed. Following careful screening, 32 publications were excluded for reasons including: focus on AI test datasets rather than software, insufficient methodological or clinical information, absence of psychiatry-specific AI applications, lack of public availability of the tools, or inaccessibility of the full text. Ultimately, 24 publications were included, yielding 24 unique psychiatry-focused AI applications for review. The Joan Briggs Critical Appraisal tool was applied to ensure that the included studies had no significant conflict of interest or methodological flaws. To assess the clinical capacity of these applications, we used the ACCS and MPCC. The ACCS provides a structured framework for measuring the difficulty of a patient case or a clinical task, allowing for an overall evaluation of clinical reasoning complexity. MPCC was used to complement this assessment, as it helps distinguish between knowledge, application, demonstration, and performance in practice. Using these different frameworks in tandem allowed us to evaluate the included studies with transparency (Joanna Briggs), capture the depth and complexity of clinical reasoning (ACCS), and assess the progression of competence from theoretical knowledge to applied clinical performance (MPCC), providing a more comprehensive evaluation of psychiatric AI tools. The included publications underwent an extensive process involving seven researchers to retrieve relevant information on AI clinical capability testing and to paraphrase key points for input in Excel spreadsheets. The retrieved information was then reviewed weekly by different researchers for two months. Any discrepancies were addressed by all researchers. There was no direct contact with any of the authors of the literature review articles. Although no statistical analysis was conducted, one researcher calculated the frequencies of peer-reviewed testing for AI clinical capability elements comparable to board certification examinations. A flowchart of the literature review search process, as described before, is shown in Figure 1. Results A total of 24 AI applications were reviewed. Based on this analysis, 58.33% of applications incorporated Generative AI technologies (including large language models), and 66.67% were specifically focused on psychiatry, with training tailored to psychiatric applications. Only 12.5% demonstrated integration of safety considerations, privacy measures, or explainability in their product design. Regarding end-user access, 79.17% allowed patients as end-users, while 50% enabled physician use. However, none of the applications that permitted physician access included access control mechanisms to restrict patient use. In terms of conflicts of interest, 54.17% of entries reported no conflicts, while 45.83% disclosed conflicts of interest. Finally, nearly all (95.83%) of the applications were accessible in North America (except Limbic AI), and 100% were available in Europe, Asia, Africa, and Latin America. Background Information After analyzing which AI applications can be used for psychiatric evaluation, we observed that 8 out of 24 applications (33.3%) meet the criteria*. Regarding which applications can be used for psychiatric diagnosis, 10 of the 14 applications (41.7%) were identified. The cross-tabulation indicates that among those rated for psychiatric evaluation, 100% of the applications marked “Yes” for evaluation also have a “Yes” for diagnosis. In contrast, the “No” evaluation group is split, with most (87.5%) not suitable for diagnosis (Figure 2). When evaluating the involvement of medical professionals in the design of AI applications, our analysis shows that only 6 of 24 AI applications (25%) involved them in the process. Clinical Capabilities by Domain Psychiatric Evaluation In the domain of psychiatric evaluation, the majority of applications (79.17%, n = 19) lacked a defined complexity rating on the ACCS. Four applications (16.67%) were classified as “moderately well defined, moderately difficult” (score 10–15), and one application (4.17%) was rated as “mildly straightforward and defined” (score 5–10). Correspondingly, 79.17% (n = 19) had no specified Miller competency rating, while 20.83% (n = 5) were categorized at Level 2, indicating the ability to apply knowledge to structured clinical cases (e.g., multiple choice questions). All applications with defined ACCS scores had Level 2 competency ratings, while those lacking ACCS scores were also marked as “none specified” on Miller’s scale (Figure 3). Psychiatric Diagnosis For psychiatric diagnosis, 70.83% of applications (n = 17) lacked a specified ACCS complexity rating. Four (16.67%) were categorized as moderately difficult (score 10–15), two (8.33%) as mildly straightforward (score 5–10), and one (4.17%) as highly complex (score 20–30). On Miller’s Pyramid, 70.83% had no competency rating, while 20.83% were at Level 2 and 8.33% at Level 4, representing real-world application of knowledge. All applications with moderate or high complexity were rated Level 2, while those with mildly straightforward ratings were rated Level 4 (Figure 4). Applications without ACCS ratings were uniformly unrated on Miller’s Pyramid. Psychopharmacology Treatment In the psychopharmacology domain, 83.33% of applications (n = 20) lacked defined complexity ratings. Three applications (12.5%) were rated as mildly straightforward, and one (4.17%) as moderately difficult. Similarly, 83.33% had no defined Miller competency level, while 16.67% were classified at Level 2. All four applications with defined ACCS scores corresponded with Level 2 competency (Figure 5). Psychotherapy Intervention The psychotherapy domain showed more diversity. Half of the applications (n = 12, 50.0%) were rated as mildly straightforward on the ACCS, while 45.83% (n = 11) received no rating, and one (4.17%) was rated as moderately difficult. On Miller’s Pyramid, 45.83% were unrated, 16.67% were classified at Level 2, and 37.5% reached Level 4. Among those with mildly straightforward complexity, 75% were rated at Level 4 and 25% at Level 2 (Figure 6). The single moderately difficult application was rated at Level 2, and all applications with undefined complexity were unrated on Miller’s scale. Psychosocial Intervention In this domain, 79.17% (n = 19) of applications lacked a defined complexity rating, while 20.83% (n = 5) were categorized as mildly straightforward. On Miller’s Pyramid, 79.17% were unrated, 16.67% were rated Level 2, and 4.17% rated Level 4. Of the five applications with defined complexity, four (80%) were rated Level 2 and one (20%) was rated Level 4 (Figure 7). Evaluation of Clinical Training, Testing, and Data Availability None of the 24 applications included in this study used formal psychiatry board certification-style examination methods as part of their internal competency testing (100%). Regarding stated clinical training, 79.17% (n = 19) of applications explicitly claimed psychiatry-specific training, while 20.83% (n = 5) did not. Among those that claimed to have psychiatric training, only 10.53% (n = 2) reported using board certification-level training material, while the remaining 89.47% (n = 17) did not. Data transparency was limited. Only one application (4.17%) provided publicly available information about the psychiatry training data used in its development; the other 23 applications (95.83%) did not disclose such information. In contrast, testing data availability was more common: 66.67% of applications (n = 16) provided at least some details on testing data, while 33.33% (n = 8) offered none. Comparative Performance of AI Applications When aggregating cumulative scores from both the ACCS and MPCC across all clinical domains, ChatGPT emerged as the top-performing application with a total score of 24.0 (ACCS: 20.0; MPCC: 4.0). Wysa followed with a score of 23.0 (ACCS: 15.0; MPCC: 8.0). DoctorGPT (ALGOMED & OpenAI), Med-PaLM, and Mellama each scored 19.0. A middle tier of applications, including CBT-i Coach, BetterHelp, TESS, and Elomia, scored 11.5. Several applications, including Talkspace, Replika, Moodpath, and PubMedBert, lacked sufficient scoring data and were assigned a cumulative score of zero. This is shown in Figure 8. Discussion Our analysis of AI applications' capacity for psychiatric evaluation and diagnosis demonstrates that any application capable of supporting psychiatric evaluation is inherently capable of aiding in diagnosis. This could be due to a shared technological or algorithmic foundation between evaluation and diagnosis functions, indicating a logical pipeline. This finding supports the argument for multi-functional design in AI applications, meaning tools should not be designed to perform only one isolated function if the underlying capability enables more. Among the applications that are not suitable for evaluation, a small fraction can support diagnosis without evaluation capabilities. This raises concerns about diagnostic robustness, and further investigation is needed to evaluate whether these applications oversimplify diagnoses. As only a third of AI applications use psychiatric evaluation, there may be a potential misalignment between clinical needs and AI app development priorities. This also indicates a gap in the AI health tech market: diagnostic tools are more prevalent, yet fewer diagnostic tools are designed with comprehensive evaluation capabilities, which are necessary for accurate, holistic, and context-aware psychiatric care. Although overall diagnostic capability data is slightly higher than evaluation, the results demonstrate that AI in psychiatry remains underdeveloped and underregulated, as many applications do not meet the criteria for either core function. Analysis of medical professional involvement in the development of diagnostic tools shows a significant finding: the majority of these AI applications for psychiatric use were developed without direct input from medical professionals. This is concerning, as it limits the clinical use of these applications and the benefit they may provide to patient populations. To build holistic computational models, clinicians' perspectives and knowledge need to be integrated with those of researchers (modellers, data scientists). With more transparency from both sides, AI applications can be more rigorously evaluated and adapted to improve clinical outcomes in the future. This study demonstrates a significant heterogeneity in the clinical capabilities of currently available psychiatric AI applications, with most failing to meet higher levels of clinical complexity or competence. Across all five core domains—psychiatric evaluation, diagnosis, psychopharmacology, psychotherapy, and psychosocial interventions—the majority of applications did not extend beyond basic case complexity as measured by the Amsterdam Clinical Challenge Scale (ACCS) or beyond foundational knowledge application on Miller’s Pyramid of Clinical Competence. Psychiatric evaluation and diagnosis, two of the most foundational domains in clinical psychiatry, were frequently underdeveloped, with nearly 80% of applications in both domains lacking defined complexity ratings or clinical reasoning capabilities. Although a small number of applications reached Miller Level 2, suggesting basic application of psychiatric knowledge, only a few achieved Level 4, which requires performance in authentic, real-world settings. These results indicate that most applications may be limited in their ability to support nuanced clinical decision-making or to adapt to the complexity of psychiatric presentations. Notably, psychotherapy was the most developed domain, with 50% of applications reaching at least mild complexity and 37.5% demonstrating Miller Level 4 competence. This may reflect the relative maturity of conversational AI platforms trained on structured psychotherapeutic modalities such as cognitive behavioral therapy (CBT). The findings also highlight the importance of supportive and behavioral interventions, which are more amenable to algorithmic structuring than diagnostic reasoning or pharmacologic decision-making. Overall, the findings suggest that current psychiatric AI tools are better suited to structured, protocol-driven interventions than to complex diagnostic or evaluative tasks requiring clinical flexibility and judgment. Our evaluation revealed significant gaps in transparency, consistency, and training standards across psychiatric AI tools. None of the applications used formal board certification-style testing to validate their clinical competence. This lack of standardized assessment raises concerns regarding the robustness and generalizability of the applications' clinical outputs, particularly in high-stakes settings. While the majority of applications (79.17%) reported some psychiatry-specific training, only a minority (10.53%) reported exposure to board certification-level material. This disparity suggests that many tools rely on general mental health or non-specialist training sources, which could limit their relevance to real-world psychiatric care. Furthermore, access to training data remains critically limited, with only one application (4.17%) providing information on its psychiatric training datasets. The absence of transparent training data precludes meaningful peer review and impedes independent validation of these tools' safety and effectiveness. In contrast, the availability of psychiatry-specific testing data was more encouraging, with 66.67% of applications providing some form of testing documentation. However, without standardization of testing procedures and outcome metrics, the interpretability and comparability of this data remain limited. Collectively, these findings suggest that most psychiatric AI tools fall short of clinical standards for training and validation. There is a clear need for regulatory frameworks that require disclosure of training sources, testing protocols, and evidence of performance on clinically relevant tasks. Among the applications evaluated, ChatGPT and Wysa emerged as the highest performers in terms of overall clinical complexity and competence. ChatGPT's strong performance may reflect its broad training data, advanced natural language processing architecture, and generalizability across domains. However, its elevated ACCS score may also reflect a higher risk of overgeneralization in psychiatric contexts, where nuance and precision are essential. Wysa’s high score, particularly in psychotherapy domains, likely stems from its focused design as a structured CBT-based tool, which aligns well with both ACCS and Miller’s criteria for clinical task execution. Other applications, such as Med-PaLM and DoctorGPT, also demonstrated promising but incomplete capabilities, often lacking depth in real-world application, despite strong knowledge representation. In contrast, a large proportion of tools either failed to provide sufficient data for scoring or lacked meaningful indicators of complexity and competence. This heterogeneity underscores the need for standardized benchmarks for evaluating psychiatric AI tools and for centralized evaluation protocols, particularly as these technologies become more integrated into clinical workflows. Conclusion Overall, our analysis demonstrates that while psychiatric AI applications show early promise, they remain limited in both scope and rigor. Most tools prioritize structured, protocol-driven interventions such as psychotherapy, while diagnostic and evaluative capacities, arguably the core of psychiatric practice, remain underdeveloped and inconsistently validated. The absence of standardized training, transparent datasets, and clinician involvement in development underscores a critical gap between technological innovation and clinical applicability. Although high-performing tools like ChatGPT and Wysa illustrate the potential of AI to augment psychiatric care, their successes highlight the need for robust regulatory frameworks, integration of medical expertise, and transparent validation protocols to ensure safety, reliability, and clinical utility. Taken together, these findings suggest that AI in psychiatry is at a pivotal stage: poised for transformative impact, but requiring deliberate design, clinician-researcher collaboration, and stronger oversight to align with the complexity of real-world psychiatric care. Declarations Competing Interests The authors declare no competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Author Contribution E.A. conceptualized the study. C.M., M.S., E.A., J.S., D.T., J.L., S.T., H.A., and A.J. contributed to data collection, analysis, and interpretation. P.C. and K.M drafted the manuscript. All authors critically revised the manuscript for important intellectual content and approved the final version. Acknowledgements The authors would like to thank the Brooklyn Brain and Mind Institute for their support during the development of this project. We are grateful to our colleagues for their contributions to literature screening and data extraction. References Esteva A, Kuprel B, Novoa RA, et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature . 2017;542(7639):115–8. Rajpurkar P, Irvin J, Zhu K, et al. CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. arXiv preprint . 2017. arXiv:1711.05225. Ahmed MI, Spooner B, Isherwood J, Lane M, Orrock E, Dennison A. A Systematic Review of the Barriers to the Implementation of Artificial Intelligence in Healthcare. Cureus. 2023;15(10):e46454. doi: 10.7759/cureus.46454 . PMID: 37927664; PMCID: PMC10623210. Lai J, Widmar NO, Liang Y. Consumers’ willingness to pay for AI doctor consultations: evidence from the US. Front Public Health . 2020;8:611152. Longoni C, Bonezzi A, Morewedge CK. Resistance to medical artificial intelligence. J Consum Res . 2019;46(4):629–50. Bzdok D, Meyer-Lindenberg A. Machine learning for precision psychiatry: opportunities and challenges. Biol Psychiatry Cogn Neurosci Neuroimaging . 2018;3(3):223–30. Mohr DC, Zhang M, Schueller SM. Personal sensing: understanding mental health using ubiquitous sensors and machine learning. Annu Rev Clin Psychol . 2017;13:23–47. Fulmer R, Joerin A, Gentile B, et al. Using psychological artificial intelligence (Tess) to relieve symptoms of depression and anxiety: randomized controlled trial. JMIR Ment Health . 2018;5(4):e64. Chekroud AM, Zotti RJ, Shehzad Z, et al. Cross-trial prediction of treatment outcome in depression: a machine learning approach. Lancet Psychiatry . 2016;3(3):243–50. Benrimoh D, Israel S, Perlman K, et al. Patient and physician perspectives on the potential of AI in psychiatry. NPJ Digit Med . 2018;1:51. Cipriani A, Geddes J, Cipriani A, et al. Artificial intelligence in mental health research: current status and future prospects. Evid Based Ment Health . 2023;26(1):1–4. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8206381","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":554290235,"identity":"b9765f88-4d2e-429e-b1da-ad61e7061e9e","order_by":0,"name":"Payton Colantonio","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDElEQVRIiWNgGAWjYFADdgjFw8DewMBMQC1jA5gCKTsA0sJzgEQtDAwSCfi1mLOfMX/wcQeDPD8z87PHHyruyPBLvjH8XFBhw8Df3p2ATYtlT45h48wzDIYzm9nMDQ6cecYjOTvHWHrGmTQGiTNnN2DTYnAgx7CZt40hweAwg5nEwbbDPAa3cwykedsOMxhI5GLXcv4NTAv7N4iWm2eMf+PVcgNuCw/Ulhs8ZnhtsZzxrHDmzDYJoF94yiTOnDnMI9mTVmbNcyaNB5dfzPmTN3z42GYjz8/evk2iouKwPT/74c23eSps5Pjbe7E7jIHDAEhJIIuBRYBxigMYMLA/QBfDFBkFo2AUjIKRDQAoj157TMRPkQAAAABJRU5ErkJggg==","orcid":"","institution":"Children's Hospital of Eastern Ontario Research Institute","correspondingAuthor":true,"prefix":"","firstName":"Payton","middleName":"","lastName":"Colantonio","suffix":""},{"id":554290236,"identity":"d159a0eb-76fd-4d65-88af-d8d12573c940","order_by":1,"name":"Katarina Milosavljevic","email":"","orcid":"","institution":"Touro College of Osteopathic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Katarina","middleName":"","lastName":"Milosavljevic","suffix":""},{"id":554290237,"identity":"79722724-2af1-4a67-8d8d-3801b2b6f9f1","order_by":2,"name":"Camille Mero","email":"","orcid":"","institution":"American University of Antigua College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Camille","middleName":"","lastName":"Mero","suffix":""},{"id":554290238,"identity":"298e867d-00ca-4ad1-8752-d015b8385a7b","order_by":3,"name":"Mrinal Sharma","email":"","orcid":"","institution":"American University of Antigua College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Mrinal","middleName":"","lastName":"Sharma","suffix":""},{"id":554290239,"identity":"dc26fdbf-3904-422c-9519-9bac0cae3ee0","order_by":4,"name":"Eleonora Achrak","email":"","orcid":"","institution":"NewYork-Presbyterian Brooklyn Methodist Hospital","correspondingAuthor":false,"prefix":"","firstName":"Eleonora","middleName":"","lastName":"Achrak","suffix":""},{"id":554290240,"identity":"144c8eaa-2c62-4d6a-9732-7f65004e4418","order_by":5,"name":"Julian Secondino","email":"","orcid":"","institution":"American University of Antigua College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Julian","middleName":"","lastName":"Secondino","suffix":""},{"id":554290241,"identity":"4ac09873-4010-457b-8cd6-a3fdcf6c4b81","order_by":6,"name":"Dominique Thurston","email":"","orcid":"","institution":"Touro College of Osteopathic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Dominique","middleName":"","lastName":"Thurston","suffix":""},{"id":554290242,"identity":"d006de98-d175-47d6-b2ba-2fd5b5b1bbbe","order_by":7,"name":"Jessica Lepore","email":"","orcid":"","institution":"Touro College of Osteopathic Medicine","correspondingAuthor":false,"prefix":"","firstName":"Jessica","middleName":"","lastName":"Lepore","suffix":""},{"id":554290243,"identity":"bcd584ca-4ee5-4e09-a645-384ad44b2524","order_by":8,"name":"Sarah Tedesco","email":"","orcid":"","institution":"Penn State Health Milton S. Hershey Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"","lastName":"Tedesco","suffix":""},{"id":554290244,"identity":"206a968d-8ac3-4e87-a3c5-9c139c4a65b0","order_by":9,"name":"Heela Azizi","email":"","orcid":"","institution":"Interfaith Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Heela","middleName":"","lastName":"Azizi","suffix":""},{"id":554290245,"identity":"b0dd62ca-e404-4240-b7ad-9d069cdaf9a4","order_by":10,"name":"Ayodeji Jolayemi","email":"","orcid":"","institution":"Interfaith Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Ayodeji","middleName":"","lastName":"Jolayemi","suffix":""}],"badges":[],"createdAt":"2025-11-25 19:53:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8206381/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8206381/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":97371159,"identity":"abb2d7dc-545c-463b-8ff1-000a33bb18a5","added_by":"auto","created_at":"2025-12-03 16:28:28","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4557412,"visible":true,"origin":"","legend":"","description":"","filename":"RevisedEvaluatingtheClinicalCompetenceofAIinPsychiatry.docx","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/ebbe33bcf57f96829672c598.docx"},{"id":97346352,"identity":"20b12e4c-5ed7-4cb5-98fc-e4c2ccf956d1","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12419,"visible":true,"origin":"","legend":"","description":"","filename":"ef9a7151e8a14e6984cc4df0edd320a0.json","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/38a425986943c1ab2e7e154b.json"},{"id":97346354,"identity":"a8f38d03-3fff-4ee8-ac51-6eb2fa205c56","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":60034,"visible":true,"origin":"","legend":"","description":"","filename":"ef9a7151e8a14e6984cc4df0edd320a01enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/d3967d9404a47e0a1da41f94.xml"},{"id":97346353,"identity":"bd0a3775-9e14-4a8e-b6bb-0d27901ebe16","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26377,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/bf4c0e44e48cca3e8621481e.png"},{"id":97346355,"identity":"2d764a9b-2836-4b37-9b82-af03086054aa","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16873,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/c2bd50bbdefc88e8211ea8c5.png"},{"id":97370128,"identity":"e747c160-9a9c-4760-bf04-1f5582432f2b","added_by":"auto","created_at":"2025-12-03 16:26:46","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":116208,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/80478a2637a94843fc67431f.png"},{"id":97346365,"identity":"424fe61b-a244-4bca-b545-c3e8b9835bde","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":126181,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/7f3904a8b77f1d5a99f07128.png"},{"id":97346373,"identity":"f37b9041-2a50-4e47-9bc9-30df5e7507e9","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":111092,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/6109811d552a820375e080a8.png"},{"id":97346371,"identity":"70461fc8-9a7e-47bd-9f0a-0aa6022f601e","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":118000,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/78630648aa207a68e48675d8.png"},{"id":97346376,"identity":"c8b09b0b-5db5-4771-9228-176cf139db11","added_by":"auto","created_at":"2025-12-03 11:48:35","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":109298,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/c6a1ca1f16969c8293c86b06.png"},{"id":97346364,"identity":"d5fe6589-77db-4943-bbc1-6a6951319dc1","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":29692,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/29074e1af34b43ff613cba86.png"},{"id":97346363,"identity":"b0c07113-a5d1-41e4-910a-56ea93162dab","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11445,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/604babbd48d1f880bad1997c.png"},{"id":97346358,"identity":"aaa7a07a-8b6c-492d-b231-5b2eafebcad2","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6259,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/10e6a5278cc3e97c54e3a968.png"},{"id":97370178,"identity":"1c800ca6-21a7-4e3d-bc79-d1e68e16913f","added_by":"auto","created_at":"2025-12-03 16:26:51","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":35703,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/32558ac124c97309130696d5.png"},{"id":97370575,"identity":"a16bde27-1ddd-4dc4-986f-4a42b0a6f1fd","added_by":"auto","created_at":"2025-12-03 16:27:37","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":37866,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/4f16d6771865f02f65e73b7a.png"},{"id":97346368,"identity":"9467a89c-046a-48f2-be1a-a6edad3b4323","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33046,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/6fb0102d5ad2d733f64c2457.png"},{"id":97346377,"identity":"b44f9418-66d8-4be5-862a-8e918ff124c5","added_by":"auto","created_at":"2025-12-03 11:48:35","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36111,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/17aac9bb6bc3ed781fcdb33e.png"},{"id":97369808,"identity":"b65064b7-4fd2-4243-a027-3cd7b4c52466","added_by":"auto","created_at":"2025-12-03 16:25:50","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33441,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/2d4cc9077e4946da05ffcafa.png"},{"id":97370924,"identity":"88edde54-f3f8-45ce-8925-85c083fd99d0","added_by":"auto","created_at":"2025-12-03 16:28:09","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":24161,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/237b2a774ab4fa1515583544.png"},{"id":97346374,"identity":"46f8ab62-3597-41b3-99f1-98339c2a6367","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"xml","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":54897,"visible":true,"origin":"","legend":"","description":"","filename":"ef9a7151e8a14e6984cc4df0edd320a01structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/ed1537bd4823386cb1a0cdee.xml"},{"id":97369699,"identity":"7247c09e-fe80-4b95-9cbc-23b2d1d83c91","added_by":"auto","created_at":"2025-12-03 16:25:36","extension":"html","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":66234,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/fd9fc4d4245fa0999d05e573.html"},{"id":97346360,"identity":"cb855b43-3406-4aa8-90e2-8ca1b22071ef","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":26377,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eStudy selection process following PRISMA guidelines.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/ca57f57d23c3a901026c9dd7.png"},{"id":97346350,"identity":"d82c0049-0fb9-4c3f-a243-6bc6ffaa1fce","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":16873,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eProportion of AI applications suitable for psychiatric evaluation and diagnosis.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/997b5e2c11ff96dad28b8994.png"},{"id":97346351,"identity":"4b60f49c-3007-4632-90f2-78c82eac29c6","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":116208,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePsychiatric Evaluation: Amsterdam Clinical Challenge Scale vs Miller’s Pyramid.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/fec34a96c31081c6ba31e963.png"},{"id":97371315,"identity":"2d759ca8-c61c-498e-94fa-948d7a92a47d","added_by":"auto","created_at":"2025-12-03 16:28:42","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":126181,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePsychiatric Diagnosis: Amsterdam Clinical Challenge Scale vs Miller’s Pyramid.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/2ecbd4a66f424881190e0568.png"},{"id":97346356,"identity":"8d01e514-d2f2-478c-bf45-eaeb26b96eaf","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":111092,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePsychopharmacology Treatment: Amsterdam Clinical Challenge Scale vs Miller’s Pyramid.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/960779da668516bc02b57fdb.png"},{"id":97371074,"identity":"3fc4ab17-5938-4403-867b-c284ae899b76","added_by":"auto","created_at":"2025-12-03 16:28:23","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":118000,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePsychotherapy Intervention: Amsterdam Clinical Challenge Scale vs Miller’s Pyramid.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/b92a37b2f1790d63b833c151.png"},{"id":97346366,"identity":"3b9f7e2e-dd5a-42ec-9980-1ac5aadc9ec3","added_by":"auto","created_at":"2025-12-03 11:48:34","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":109298,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePsychosocial Intervention: Amsterdam Clinical Challenge Scale vs Miller’s Pyramid.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/38533aa820e13d10a8e085e5.png"},{"id":97371223,"identity":"13f95ba1-59bf-4627-b750-e49d86cc8a58","added_by":"auto","created_at":"2025-12-03 16:28:33","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":29692,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eRanking of applications based on cumulative Amsterdam and Miller scores.\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/57fa39e24b7d283917e22765.png"},{"id":99787890,"identity":"2ad9e1de-b1dc-4d8d-984c-2464942f56ba","added_by":"auto","created_at":"2026-01-08 12:40:50","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1426476,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8206381/v1/9f5f77fc-9e57-48c5-ad4c-d38419ae3946.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluating the Clinical Competence of Artificial Intelligence Applications in Psychiatry: A Systematic Review","fulltext":[{"header":"Introduction","content":"\u003cp\u003eArtificial Intelligence (AI) refers to computational systems capable of performing tasks that traditionally require human intelligence, such as pattern recognition, problem-solving, and decision-making. In healthcare, AI applications have rapidly expanded over the last decade due to advances in machine learning algorithms, natural language processing, and increased computational power. AI has demonstrated utility across diverse specialties, including radiology, oncology, dermatology, and cardiology, where it has been applied to diagnostic imaging, treatment planning, and risk stratification, with performance at times comparable to or exceeding that of human clinicians\u003csup\u003e1,2\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eIn psychiatry, AI offers a unique promise given the discipline’s reliance on subjective assessments, longitudinal monitoring, and multimodal data integration. Within psychiatric practice, AI encompasses tools such as natural language processing models for clinical documentation and diagnostic support, conversational agents for delivering psychotherapy, predictive algorithms for relapse or suicide risk, and decision support systems for psychopharmacological treatment. Examples include smartphone-based monitoring of mood and behavior, virtual therapeutic chatbots, and large language models (LLMs) capable of synthesizing diagnostic criteria from structured and unstructured data.\u003c/p\u003e\n\u003cp\u003eDespite this potential, the adoption of AI in psychiatry remains limited. Systematic reviews of healthcare AI have identified several barriers, including insufficient data transparency, a lack of standardized evaluation frameworks, and limited clinician involvement in the design and validation processes. Unlike radiology or pathology, psychiatry lacks large standardized datasets with objective biomarkers, relying instead on nuanced clinical interviews, patient self-reports, and complex social contexts that challenge algorithmic modelling. Clinicians have expressed skepticism regarding whether current AI systems can adequately capture the subtleties of psychiatric assessment, particularly in culturally diverse populations and comorbid conditions. Furthermore, patients may be hesitant to accept diagnoses or treatment recommendations from AI systems, particularly when such recommendations are generated through opaque, “black box” processes that lack explainability. These challenges highlight why, compared with other medical specialties, psychiatry has been slower to integrate AI technologies, and raise the question of whether existing tools are sufficiently ready for real-world use to function safely and competently in clinical environments.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIf successfully integrated, however, AI applications have the potential to transform psychiatric care by enhancing diagnostic accuracy, personalizing treatment, and expanding access to care. For example, AI-driven digital phenotyping has been explored to predict depressive episodes based on smartphone use patterns, while conversational agents delivering cognitive behavioral therapy (CBT) have shown efficacy in reducing symptoms of anxiety and depression\u003csup\u003e7,8\u003c/sup\u003e. Predictive algorithms are also being developed to forecast treatment response trajectories and inform psychopharmacological decisions\u003csup\u003e9\u003c/sup\u003e. These applications suggest that AI could serve as a valuable adjunct to psychiatrists, augmenting, rather than replacing, clinical judgment.\u003c/p\u003e\n\u003cp\u003eThe challenges of integrating AI into psychiatric care, despite the promises, are motivated by clinician skepticism. Current AI systems in psychiatry often lack rigorous validation against gold-standard clinical benchmarks, such as board certification-style examinations or structured clinical assessments. The current methods of testing LLM based AI applications for healthcare, for instance, include a variety of question banks drawn from medical student examinations. It is unclear whether AI applications could provide meaningful clinical decisions for real-world clinical scenarios. This lack of real-world, realistic competency testing undermines confidence that AI systems are truly ready for clinical use. Moreover, many tools are developed without meaningful input from psychiatrists during design, limiting their clinical applicability\u003csup\u003e10\u003c/sup\u003e. Issues of transparency in training datasets, heterogeneity in performance across psychiatric domains, and the lack of standardized evaluation protocols remain substantial barriers to trust and implementation\u003csup\u003e11\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eIn this study, we present a systematic review of AI applications in psychiatry, with a focus on tests for clinical competence. We rate the tests used by the AI applications for clinical complexity, as well as the sophistication of clinical decisions by these applications. Standardized ratings for clinical complexity (Amsterdam Clinical Challenge Scale) and clinical decision sophistication (Miller’s Pyramid of Clinical Competence) were used. The rating allowed for a more uniform comparison between the various AI applications tested with various question banks. The rating also allowed for comparison against human physicians who are certified using board certification examinations.\u003c/p\u003e\n\u003cp\u003eBy systematically analyzing the peer-reviewed literature, this review aims to address clinical skepticism and characterize the current state of AI applications in psychiatry in terms of clinical competence. This is the first known review that seeks to provide a more rigorous understanding of the clinical competence of AI applications in psychiatry, as well as to rank these applications against humans based on their ability to address real-world clinical scenarios. We also seek to highlight gaps in clinical testing and identify areas where rigorous validation is most urgently needed. Understanding these limitations and opportunities is essential to guide future development of AI systems that can achieve true real-world readiness and safely support psychiatric care.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eA total of 24 AI applications were evaluated for their clinical capabilities in psychiatry across five core domains: psychiatric evaluation, psychiatric diagnosis, psychopharmacology, psychotherapy, and psychosocial interventions. Each application was assessed using Joan’s Briggs Critical Appraisal Tool to ensure transparency, \u0026nbsp; Amsterdam Clinical Challenge Scale (ACCS) to evaluate case complexity, and Miller’s Pyramid of Clinical Competence (MPCC) to determine the depth of clinical reasoning and applied knowledge.\u003c/p\u003e\n\u003cp\u003eWe conducted a literature review from March 2024 to July 2025, focusing on peer-reviewed publications that evaluate the clinical capabilities of AI software in Psychiatry. The MOOSE guidelines for Meta-Analysis and Systematic Reviews for Observational Studies were followed to identify AI software accurately. Literature reviews were conducted using PubMed/MEDLINE, Google Scholar, Ovid, JoVE, and ScienceDirect. The research was conducted using search terms including but not limited to “artificial intelligence,” “AI”, “Psychiatry”, “testing”, “board certification”, “examinations”, “Clinical Skills”, “clinical decision”, “clinical capabilities”, “Clinical Competency,” “Performance”, “human comparison”. Search terms also included publicly available Psychiatry AI softwares including but not limited to \"ChatGPT DSM V Focus4\", \"Med-PaLM\", \"Dragon\", \"AMIE\", \"ChatGPT DSM-5 Focused Assistant\", \"BiolinkBert\", \"DrugCard\", \"PubMedGPT\", \"MindMateGPT\", \"YourDoctor\", \"PubMedBert\", \"Dedicated Psychiatrist\", \"Psyscribe\", \"Woebot\", \"DoctorGPT ALGOMED\", \"RxGuide (ChatGPT)\", \"Doctor GPT Chat GPT\", \"Prognos Application\", \"Soap note AI\", \"Ada Health\", \"Wysa\", \"Replika\", \"Talkspace\", \"Youper\", \"BetterHelp\", \"Moodpath aka MindDoc\", \"Kintsugi\", \"ReGain\", \"SuperBetter\", \"CBT-i Coach\", \"Happify\", \"MindShift\", \"MoodMission\", \"Headspace\", \"Calm\", \"Daylio\", \"RDMaster\", \"Limbic AI\", \"Moodpen\", \"Calm Mind\", \"Mind well AI\", \"Aiberry\", \"MyClinicalWriter\", \"Ree AI\", \"Corti\", \"Elomia\", \"Meomind\", Sintelly\", \"DBT Diary Card and Skills Coach\", \"MoodTools\" and \"Calm Harm\". \u0026nbsp;These search terms were used in all possible combinations.\u003c/p\u003e\n\u003cp\u003ePapers were included if they were published after April 2014, included any of the searched Psychiatry AI software as a subject, and the AI software was publicly available to patients or clinicians. All papers were eligible for inclusion if they included tests on the clinical capabilities of the Psychiatry AI software. Papers involving AI software that were not publicly available for use were excluded from the study. The end users of the AI software (physicians or patients) were not an exclusion criterion in the literature review. The type of AI model used in the AI software and its scope of application were considered, but did not alter the inclusion criteria.*\u003c/p\u003e\n\u003cp\u003eA total of 68 papers were identified through database searching, of which 56 publications remained after duplicates were removed. Following careful screening, 32 publications were excluded for reasons including: focus on AI test datasets rather than software, insufficient methodological or clinical information, absence of psychiatry-specific AI applications, lack of public availability of the tools, or inaccessibility of the full text. Ultimately, 24 publications were included, yielding 24 unique psychiatry-focused AI applications for review. The Joan Briggs Critical Appraisal tool was applied to ensure that the included studies had no significant conflict of interest or methodological flaws. To assess the clinical capacity of these applications, we used the ACCS and MPCC. The ACCS provides a structured framework for measuring the difficulty of a patient case or a clinical task, allowing for an overall evaluation of clinical reasoning complexity. MPCC was used to complement this assessment, as it helps distinguish between knowledge, application, demonstration, and performance in practice. Using these different frameworks in tandem allowed us to evaluate the included studies with transparency (Joanna Briggs), capture the depth and complexity of clinical reasoning (ACCS), and assess the progression of competence from theoretical knowledge to applied clinical performance (MPCC), providing a more comprehensive evaluation of psychiatric AI tools.\u003c/p\u003e\n\u003cp\u003eThe included publications underwent an extensive process involving seven researchers to retrieve relevant information on AI clinical capability testing and to paraphrase key points for input in Excel spreadsheets. The retrieved information was then reviewed weekly by different researchers for two months. Any discrepancies were addressed by all researchers. There was no direct contact with any of the authors of the literature review articles. Although no statistical analysis was conducted, one researcher calculated the frequencies of peer-reviewed testing for AI clinical capability elements comparable to board certification examinations.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;A flowchart of the literature review search process, as described before, is shown in Figure 1.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eA total of 24 AI applications were reviewed. Based on this analysis, 58.33% of applications incorporated Generative AI technologies (including large language models), and 66.67% were specifically focused on psychiatry, with training tailored to psychiatric applications. Only 12.5% demonstrated integration of safety considerations, privacy measures, or explainability in their product design. Regarding end-user access, 79.17% allowed patients as end-users, while 50% enabled physician use. However, none of the applications that permitted physician access included access control mechanisms to restrict patient use. In terms of conflicts of interest, 54.17% of entries reported no conflicts, while 45.83% disclosed conflicts of interest. Finally, nearly all (95.83%) of the applications were accessible in North America (except Limbic AI), and 100% were available in Europe, Asia, Africa, and Latin America.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBackground Information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter analyzing which AI applications can be used for psychiatric evaluation, we observed that 8 out of 24 applications (33.3%) meet the criteria*. Regarding which applications can be used for psychiatric diagnosis, 10 of the 14 applications (41.7%) were identified. The cross-tabulation indicates that among those rated for psychiatric evaluation, 100% of the applications marked \u0026ldquo;Yes\u0026rdquo; for evaluation also have a \u0026ldquo;Yes\u0026rdquo; for diagnosis. In contrast, the \u0026ldquo;No\u0026rdquo; evaluation group is split, with most (87.5%) not suitable for diagnosis (Figure 2).\u003c/p\u003e\n\u003cp\u003eWhen evaluating the involvement of medical professionals in the design of AI applications, our analysis shows that only 6 of 24 AI applications (25%) involved them in the process.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eClinical Capabilities by Domain\u003c/u\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePsychiatric Evaluation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the domain of psychiatric evaluation, the majority of applications (79.17%, n = 19) lacked a defined complexity rating on the ACCS. Four applications (16.67%) were classified as \u0026ldquo;moderately well defined, moderately difficult\u0026rdquo; (score 10\u0026ndash;15), and one application (4.17%) was rated as \u0026ldquo;mildly straightforward and defined\u0026rdquo; (score 5\u0026ndash;10). Correspondingly, 79.17% (n = 19) had no specified Miller competency rating, while 20.83% (n = 5) were categorized at Level 2, indicating the ability to apply knowledge to structured clinical cases (e.g., multiple choice questions). All applications with defined ACCS scores had Level 2 competency ratings, while those lacking ACCS scores were also marked as \u0026ldquo;none specified\u0026rdquo; on Miller\u0026rsquo;s scale (Figure 3).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePsychiatric Diagnosis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor psychiatric diagnosis, 70.83% of applications (n = 17) lacked a specified ACCS complexity rating. Four (16.67%) were categorized as moderately difficult (score 10\u0026ndash;15), two (8.33%) as mildly straightforward (score 5\u0026ndash;10), and one (4.17%) as highly complex (score 20\u0026ndash;30). On Miller\u0026rsquo;s Pyramid, 70.83% had no competency rating, while 20.83% were at Level 2 and 8.33% at Level 4, representing real-world application of knowledge. All applications with moderate or high complexity were rated Level 2, while those with mildly straightforward ratings were rated Level 4 (Figure 4). Applications without ACCS ratings were uniformly unrated on Miller\u0026rsquo;s Pyramid.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePsychopharmacology Treatment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the psychopharmacology domain, 83.33% of applications (n = 20) lacked defined complexity ratings. Three applications (12.5%) were rated as mildly straightforward, and one (4.17%) as moderately difficult. Similarly, 83.33% had no defined Miller competency level, while 16.67% were classified at Level 2. All four applications with defined ACCS scores corresponded with Level 2 competency (Figure 5).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePsychotherapy Intervention\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe psychotherapy domain showed more diversity. Half of the applications (n = 12, 50.0%) were rated as mildly straightforward on the ACCS, while 45.83% (n = 11) received no rating, and one (4.17%) was rated as moderately difficult. On Miller\u0026rsquo;s Pyramid, 45.83% were unrated, 16.67% were classified at Level 2, and 37.5% reached Level 4. Among those with mildly straightforward complexity, 75% were rated at Level 4 and 25% at Level 2 (Figure 6). The single moderately difficult application was rated at Level 2, and all applications with undefined complexity were unrated on Miller\u0026rsquo;s scale.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePsychosocial Intervention\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this domain, 79.17% (n = 19) of applications lacked a defined complexity rating, while 20.83% (n = 5) were categorized as mildly straightforward. On Miller\u0026rsquo;s Pyramid, 79.17% were unrated, 16.67% were rated Level 2, and 4.17% rated Level 4. Of the five applications with defined complexity, four (80%) were rated Level 2 and one (20%) was rated Level 4 (Figure 7).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEvaluation of Clinical Training, Testing, and Data Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone of the 24 applications included in this study used formal psychiatry board certification-style examination methods as part of their internal competency testing (100%).\u003c/p\u003e\n\u003cp\u003eRegarding stated clinical training, 79.17% (n = 19) of applications explicitly claimed psychiatry-specific training, while 20.83% (n = 5) did not. Among those that claimed to have psychiatric training, only 10.53% (n = 2) reported using board certification-level training material, while the remaining 89.47% (n = 17) did not.\u003c/p\u003e\n\u003cp\u003eData transparency was limited. Only one application (4.17%) provided publicly available information about the psychiatry training data used in its development; the other 23 applications (95.83%) did not disclose such information. In contrast, testing data availability was more common: 66.67% of applications (n = 16) provided at least some details on testing data, while 33.33% (n = 8) offered none.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComparative Performance of AI Applications\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhen aggregating cumulative scores from both the ACCS and MPCC across all clinical domains, ChatGPT emerged as the top-performing application with a total score of 24.0 (ACCS: 20.0; MPCC: 4.0). Wysa followed with a score of 23.0 (ACCS: 15.0; MPCC: 8.0). DoctorGPT (ALGOMED \u0026amp; OpenAI), Med-PaLM, and Mellama each scored 19.0. A middle tier of applications, including CBT-i Coach, BetterHelp, TESS, and Elomia, scored 11.5. Several applications, including Talkspace, Replika, Moodpath, and PubMedBert, lacked sufficient scoring data and were assigned a cumulative score of zero. This is shown in Figure 8.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur analysis of AI applications' capacity for psychiatric evaluation and diagnosis demonstrates that any application capable of supporting psychiatric evaluation is inherently capable of aiding in diagnosis. This could be due to a shared technological or algorithmic foundation between evaluation and diagnosis functions, indicating a logical pipeline. This finding supports the argument for multi-functional design in AI applications, meaning tools should not be designed to perform only one isolated function if the underlying capability enables more. Among the applications that are not suitable for evaluation, a small fraction can support diagnosis without evaluation capabilities. This raises concerns about diagnostic robustness, and further investigation is needed to evaluate whether these applications oversimplify diagnoses. As only a third of AI applications use psychiatric evaluation, there may be a potential misalignment between clinical needs and AI app development priorities. This also indicates a gap in the AI health tech market: diagnostic tools are more prevalent, yet fewer diagnostic tools are designed with comprehensive evaluation capabilities, which are necessary for accurate, holistic, and context-aware psychiatric care. Although overall diagnostic capability data is slightly higher than evaluation, the results demonstrate that AI in psychiatry remains underdeveloped and underregulated, as many applications do not meet the criteria for either core function.\u003c/p\u003e\u003cp\u003eAnalysis of medical professional involvement in the development of diagnostic tools shows a significant finding: the majority of these AI applications for psychiatric use were developed without direct input from medical professionals. This is concerning, as it limits the clinical use of these applications and the benefit they may provide to patient populations. To build holistic computational models, clinicians' perspectives and knowledge need to be integrated with those of researchers (modellers, data scientists). With more transparency from both sides, AI applications can be more rigorously evaluated and adapted to improve clinical outcomes in the future.\u003c/p\u003e\u003cp\u003eThis study demonstrates a significant heterogeneity in the clinical capabilities of currently available psychiatric AI applications, with most failing to meet higher levels of clinical complexity or competence. Across all five core domains\u0026mdash;psychiatric evaluation, diagnosis, psychopharmacology, psychotherapy, and psychosocial interventions\u0026mdash;the majority of applications did not extend beyond basic case complexity as measured by the Amsterdam Clinical Challenge Scale (ACCS) or beyond foundational knowledge application on Miller\u0026rsquo;s Pyramid of Clinical Competence.\u003c/p\u003e\u003cp\u003ePsychiatric evaluation and diagnosis, two of the most foundational domains in clinical psychiatry, were frequently underdeveloped, with nearly 80% of applications in both domains lacking defined complexity ratings or clinical reasoning capabilities. Although a small number of applications reached Miller Level 2, suggesting basic application of psychiatric knowledge, only a few achieved Level 4, which requires performance in authentic, real-world settings. These results indicate that most applications may be limited in their ability to support nuanced clinical decision-making or to adapt to the complexity of psychiatric presentations.\u003c/p\u003e\u003cp\u003eNotably, psychotherapy was the most developed domain, with 50% of applications reaching at least mild complexity and 37.5% demonstrating Miller Level 4 competence. This may reflect the relative maturity of conversational AI platforms trained on structured psychotherapeutic modalities such as cognitive behavioral therapy (CBT). The findings also highlight the importance of supportive and behavioral interventions, which are more amenable to algorithmic structuring than diagnostic reasoning or pharmacologic decision-making.\u003c/p\u003e\u003cp\u003eOverall, the findings suggest that current psychiatric AI tools are better suited to structured, protocol-driven interventions than to complex diagnostic or evaluative tasks requiring clinical flexibility and judgment.\u003c/p\u003e\u003cp\u003eOur evaluation revealed significant gaps in transparency, consistency, and training standards across psychiatric AI tools. None of the applications used formal board certification-style testing to validate their clinical competence. This lack of standardized assessment raises concerns regarding the robustness and generalizability of the applications' clinical outputs, particularly in high-stakes settings.\u003c/p\u003e\u003cp\u003eWhile the majority of applications (79.17%) reported some psychiatry-specific training, only a minority (10.53%) reported exposure to board certification-level material. This disparity suggests that many tools rely on general mental health or non-specialist training sources, which could limit their relevance to real-world psychiatric care. Furthermore, access to training data remains critically limited, with only one application (4.17%) providing information on its psychiatric training datasets. The absence of transparent training data precludes meaningful peer review and impedes independent validation of these tools' safety and effectiveness.\u003c/p\u003e\u003cp\u003eIn contrast, the availability of psychiatry-specific testing data was more encouraging, with 66.67% of applications providing some form of testing documentation. However, without standardization of testing procedures and outcome metrics, the interpretability and comparability of this data remain limited.\u003c/p\u003e\u003cp\u003eCollectively, these findings suggest that most psychiatric AI tools fall short of clinical standards for training and validation. There is a clear need for regulatory frameworks that require disclosure of training sources, testing protocols, and evidence of performance on clinically relevant tasks.\u003c/p\u003e\u003cp\u003eAmong the applications evaluated, ChatGPT and Wysa emerged as the highest performers in terms of overall clinical complexity and competence. ChatGPT's strong performance may reflect its broad training data, advanced natural language processing architecture, and generalizability across domains. However, its elevated ACCS score may also reflect a higher risk of overgeneralization in psychiatric contexts, where nuance and precision are essential.\u003c/p\u003e\u003cp\u003eWysa\u0026rsquo;s high score, particularly in psychotherapy domains, likely stems from its focused design as a structured CBT-based tool, which aligns well with both ACCS and Miller\u0026rsquo;s criteria for clinical task execution. Other applications, such as Med-PaLM and DoctorGPT, also demonstrated promising but incomplete capabilities, often lacking depth in real-world application, despite strong knowledge representation.\u003c/p\u003e\u003cp\u003eIn contrast, a large proportion of tools either failed to provide sufficient data for scoring or lacked meaningful indicators of complexity and competence. This heterogeneity underscores the need for standardized benchmarks for evaluating psychiatric AI tools and for centralized evaluation protocols, particularly as these technologies become more integrated into clinical workflows.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eOverall, our analysis demonstrates that while psychiatric AI applications show early promise, they remain limited in both scope and rigor. Most tools prioritize structured, protocol-driven interventions such as psychotherapy, while diagnostic and evaluative capacities, arguably the core of psychiatric practice, remain underdeveloped and inconsistently validated. The absence of standardized training, transparent datasets, and clinician involvement in development underscores a critical gap between technological innovation and clinical applicability. Although high-performing tools like ChatGPT and Wysa illustrate the potential of AI to augment psychiatric care, their successes highlight the need for robust regulatory frameworks, integration of medical expertise, and transparent validation protocols to ensure safety, reliability, and clinical utility. Taken together, these findings suggest that AI in psychiatry is at a pivotal stage: poised for transformative impact, but requiring deliberate design, clinician-researcher collaboration, and stronger oversight to align with the complexity of real-world psychiatric care.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003ch2\u003eCompeting Interests\u003c/h2\u003e\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eE.A. conceptualized the study. C.M., M.S., E.A., J.S., D.T., J.L., S.T., H.A., and A.J. contributed to data collection, analysis, and interpretation. P.C. and K.M drafted the manuscript. All authors critically revised the manuscript for important intellectual content and approved the final version.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e\u003cp\u003eThe authors would like to thank the Brooklyn Brain and Mind Institute for their support during the development of this project. We are grateful to our colleagues for their contributions to literature screening and data extraction.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eEsteva A, Kuprel B, Novoa RA, et al. Dermatologist-level classification of skin cancer with deep neural networks. \u003cem\u003eNature\u003c/em\u003e. 2017;542(7639):115\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRajpurkar P, Irvin J, Zhu K, et al. CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. \u003cem\u003earXiv preprint\u003c/em\u003e. 2017. arXiv:1711.05225.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAhmed MI, Spooner B, Isherwood J, Lane M, Orrock E, Dennison A. A Systematic Review of the Barriers to the Implementation of Artificial Intelligence in Healthcare. Cureus. 2023;15(10):e46454. doi: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.7759/cureus.46454\u003c/span\u003e\u003cspan address=\"10.7759/cureus.46454\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 37927664; PMCID: PMC10623210.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLai J, Widmar NO, Liang Y. Consumers\u0026rsquo; willingness to pay for AI doctor consultations: evidence from the US. \u003cem\u003eFront Public Health\u003c/em\u003e. 2020;8:611152.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLongoni C, Bonezzi A, Morewedge CK. Resistance to medical artificial intelligence. \u003cem\u003eJ Consum Res\u003c/em\u003e. 2019;46(4):629\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBzdok D, Meyer-Lindenberg A. Machine learning for precision psychiatry: opportunities and challenges. \u003cem\u003eBiol Psychiatry Cogn Neurosci Neuroimaging\u003c/em\u003e. 2018;3(3):223\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMohr DC, Zhang M, Schueller SM. Personal sensing: understanding mental health using ubiquitous sensors and machine learning. \u003cem\u003eAnnu Rev Clin Psychol\u003c/em\u003e. 2017;13:23\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFulmer R, Joerin A, Gentile B, et al. Using psychological artificial intelligence (Tess) to relieve symptoms of depression and anxiety: randomized controlled trial. \u003cem\u003eJMIR Ment Health\u003c/em\u003e. 2018;5(4):e64.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChekroud AM, Zotti RJ, Shehzad Z, et al. Cross-trial prediction of treatment outcome in depression: a machine learning approach. \u003cem\u003eLancet Psychiatry\u003c/em\u003e. 2016;3(3):243\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBenrimoh D, Israel S, Perlman K, et al. Patient and physician perspectives on the potential of AI in psychiatry. \u003cem\u003eNPJ Digit Med\u003c/em\u003e. 2018;1:51.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCipriani A, Geddes J, Cipriani A, et al. Artificial intelligence in mental health research: current status and future prospects. \u003cem\u003eEvid Based Ment Health\u003c/em\u003e. 2023;26(1):1\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8206381/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8206381/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eArtificial intelligence (AI) is an increasingly promising technology in psychiatry with the potential to transform mental healthcare. The number of AI applications designed for psychiatric diagnosis and treatment has grown substantially over the last few years; however, clinicians remain concerned about the real-world readiness of these applications in clinical care settings. These concerns have some validity, given recently reported cases of worsening psychosis and suicide attempts associated with AI application use. While there are a few studies that have examined the performance of AI applications on various multiple-choice question banks, the clinical usefulness and practical application of such test performance have yet to be assessed. Assessing test performance will provide an appraisal of the case complexity and appropriateness of the clinical decision-making process used by AI applications. Such an assessment offers a more definitive understanding of the real-world readiness of various AI applications in clinical scenarios, as well as their comparative performance relative to that of human physicians.\u003c/p\u003e\u003cp\u003e We conducted a systematic review of peer-reviewed publications from April 2014 to July 2025, evaluating the real-world clinical readiness of publicly available psychiatric AI applications. Following PRISMA and MOOSE guidelines, 68 publications were identified, of which 24 met the inclusion criteria, yielding 24 unique applications. Each AI application was evaluated based on the test administered for clinical competence, with case complexity rated using the Amsterdam Clinical Challenge Scale and clinical decision-making assessed via Miller\u0026rsquo;s Pyramid of Clinical Competence. The inclusion of the various domains of psychiatric practice\u0026mdash;evaluation, diagnosis, psychopharmacology, psychotherapy, and psychosocial intervention was also taken into consideration. Competence was determined based on the intersection of case complexity and decision-making level.\u003c/p\u003e\u003cp\u003eFindings revealed substantial heterogeneity in performance across the various domains of psychiatry, with no AI application achieving human-level competence in all domains of psychiatric practice. Some psychotherapy-focused tools demonstrated moderate complexity and comparatively higher competence, while most applications remained limited to basic knowledge application (Miller Level 2). Core clinical domains, such as evaluation and diagnosis, were underdeveloped, with more than 70% of tools lacking defined complexity or competence ratings. Importantly, none of the applications used board certification-style testing; only 10.5% reported training on board-level material, and just 4.2% disclosed the use of training datasets.\u003c/p\u003e\u003cp\u003eThese results highlight a critical gap between the potential of psychiatric AI and its demonstrated real-world readiness. Standardized, board-style evaluation, transparency in training data, and more rigorous measures of clinical decision-making are essential to building trust and supporting safe integration into psychiatric care.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e","manuscriptTitle":"Evaluating the Clinical Competence of Artificial Intelligence Applications in Psychiatry: A Systematic Review","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-03 11:48:29","doi":"10.21203/rs.3.rs-8206381/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d856b910-17d5-4043-a5a6-116178c47411","owner":[],"postedDate":"December 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":58987485,"name":"Scientific community and society/Business and industry"},{"id":58987486,"name":"Health sciences/Diseases"},{"id":58987487,"name":"Health sciences/Health care"},{"id":58987488,"name":"Health sciences/Medical research"},{"id":58987489,"name":"Biological sciences/Psychology"},{"id":58987490,"name":"Social science/Psychology"}],"tags":[],"updatedAt":"2025-12-23T23:08:46+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-03 11:48:29","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8206381","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8206381","identity":"rs-8206381","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00