Clinicians’ Practices and Attitudes Toward Depression Assessment: A Cross-Sectional Survey Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Clinicians’ Practices and Attitudes Toward Depression Assessment: A Cross-Sectional Survey Study Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8951240/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Accurate psychological evaluation is crucial for early detection, intervention and evaluation of mental health disorders. The validity and reliability of psychological evaluation methods have been systematically researched for many decades. However, assessment methods used by mental health professionals in clinical settings have rarely been empirically documented. This study examines the extent to which empirically supported assessments are used in clinical settings and clinicians’ attitudes toward these practices. Clinicians ( N = 585) from the U.S. (25%), the U.K. (28%), the Netherlands (18%), Sweden (22%), and other countries (7%) completed an online survey reporting their depression assessment methods. Clinicians reported using many different data collection methods (e.g., unstructured clinical interviews [83%], rating scales [83%]) in various combinations. A majority (80%) primarily employed clinical methods, as opposed to statistical data combination methods (11%). Although they were confident in their assessments, the time to carry out assessments varied widely ( M = 87 [ SD = 108] minutes). Seventy-seven percent agreed on the importance of standardized, valid, and reliable assessments, and 43% reported interest in receiving AI decision-support. We did not find many significant cross-cultural differences. A majority of clinicians have not fully adopted standardized empirically supported assessment practices. Predominant reliance on clinical judgment suggests a disconnect between best practices supported by research and everyday clinical practice. clinical assessment depression structured methods statistical models artificial intelligence Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Accurate assessment of mental health disorders is crucial for early detection and effective treatment planning. Inaccurate assessment of the type and severity of mental disorders can have significant negative impacts on people’s lives, highlighting the importance of precise assessment methods for effective prevention, early detection, and appropriate treatment. 1 – 3 There are many different methods used to assess mental health problems, including various data collection methods such as interviews and rating scales, as well as various data combination methods , i.e., approaches for evaluating, integrating, and interpreting different sources of information (Fig. 1 ). 4 , 5 The validity and reliability of depression assessment methods have been extensively studied, 6,7 with some methods being more reliable and valid than others are. However, less is known about the practices of mental health professionals, and the available evidence suggests a discrepancy between guidelines on standards of practice and those on clinical practice. 8 This study aims to gain insights into the extent to which empirically supported assessments are used by mental health professionals, as well as their attitudes toward these methods, which is critical for identifying risks and challenges in current assessment practices and the implementation of new assessment methods. Our study may also add to the discourse on integrating artificial intelligence-based (AI) decision-support systems into assessment processes. 9 – 11 Note The data combination method does not necessarily require data from different data collection methods; for example, responses to items of a single rating scale can be aggregated and interpreted via a statistical model (following an aggregation schema or using an established cut-off point) or interpreted via a clinical method in which individual items are examined via clinical experience. In the statistical method, it is also feasible to include data from a clinical (unstructured) interview in a statistical model (e.g., to what extent the patient responds, their response times, and behaviors). Data collection methods Data collection methods range from unstructured to structured . 12 Unstructured data collection methods include unstructured clinical interviews , which allow clinicians to explore patients’ issues freely and adjust the direction of the assessment on the basis of patient’s responses and behaviors without predetermined questions. Structured data collection methods are highly organized and follow a predetermined format, ensuring that the same questions are asked of each person. Structured methods aim to yield more reliable and comparable results and are especially well suited for assessing a set of validated criteria for a disorder or clinical phenomenon (e.g., evaluating whether a patient meets a diagnostic manual's criteria for a major depressive episode). Examples include rating scales and structured clinical interviews, as well as semistructured clinical interviews where questions are predetermined, but the interviewer is allowed to ask additional questions or exclude ones to adjust the sourcing of information on a case-to-case basis. The interviews were conducted by clinicians, while for the rating scales, some were clinician-rated, and others were self-reported. Data combination methods Two primary approaches for evaluating, integrating, and interpreting different sources of information have been identified for depression assessment: the clinical approach and the statistical approach. 6 , 13 , 14 The clinical data combination method is based on human judgment. Clinicians assess various sources of information on the basis of personal judgment, clinical experience, theoretical perspectives as well as surrounding and individual factors (e.g., subtle deviations from a patient’s baseline behavior, appearance, overt signs such as direct statements of distress or crying, or general deduction on the basis of the overall information gathered). In contrast, the statistical data combination method makes use of a statistical model that is based on empirically established relationships between the different sources of information and the outcome of interest. This can, for example, involve the clinician entering client data into formulas, actuarial tables, or charts. 13 , 14 To establish relationships among variables, statistical models may be constructed via relatively simple statistical approaches such as linear or logistic regression. It can also involve more complex statistical techniques, such as AI, which includes techniques such as machine learning and natural language processing. Statistical data combination methods can, for example, be used for screening, comprehensive diagnostic assessment support, treatment personalization, or symptom monitoring. 15 , 16 A growing body of research supports the use of AI in mental health care, demonstrating its ability to improve the accuracy and efficiency of diagnostic assessments. 10 , 16 , 17 Given its promising applications and increasing capabilities, we focus on AI as a key component of the statistical approach in the following sections. Final assessment and decision-making In the data combination phase, both the clinical and the statistical data combination methods can be used at different stages (see Brunswick Lens Model). 18 , 19 , 20 For example, a clinician may use a statistical method to integrate a patient’s rating scale responses but rely on a clinical method to combine those results with information from an unstructured interview. However, when the results from the two different methods disagree, one cannot follow both of them: When they disagree , the final decision will necessarily have to rely on one method over the other. 14 , 21 Navigating and integrating different information via the clinical data combination method can be cognitively strenuous and complex, 22,23 especially when the various sources are in conflict with one another. The clinical method has been shown to be less reliable than statistical models. 13 , 21 However, the implementation of statistical models can be difficult due to insufficient guidelines, education, or resources, such as a lack of validated advanced statistical models available for practical use in psychological assessments. 11 For example, models that enable the statistical combination of data from several rating scales and clinical history are still uncommon. Since the introduction of the clinical versus statistical controversy in the 1950s, statistical modeling has undergone significant improvements as a result of advancements in the field of AI. 9,10,24 Statistical models, often including AI technology, have become increasingly prevalent in healthcare, especially in medicine. 25 , 26 Nevertheless, the question remains whether the benefits of such instruments are used for depression assessments in clinical practice and, if so, to what extent, in what way, and how clinicians’ attitudes reflect their tendencies to use the statistical data combination methods. Research support for assessment methods The validity and reliability of depression assessment methods have been systematically researched for many decades. Semistructured or structured interviews are often regarded as the benchmark standard for depression assessment because they demonstrate greater validity and reliability than unstructured interviews do. 7,27,28 In contrast, unstructured clinical interviews, although widely used in practice, are less reliable and less valid, 12,29 and they are rarely used in research. Self-report questionnaires, while highly reliable, have other shortcomings. For example, rating scales targeting depression: i) fail to reliably capture the disorder across scales, 30 ii) miss important symptoms, 31 iii) differ substantially in their categorization of depressed patients into severity groups, 32 and iv) are multidimensional, thus multiple constructs with inconsistent factor structures across scales are assessed. 33 – 35 Additionally, the statistical method for combining information, which involves the use of algorithms or structured approaches, generally results in a 13% increase in accuracy over the clinical method, which relies on the clinician's judgment. 6 , 13 , 24 , 36 Despite the strong evidence supporting the use of structured assessments and statistical methods, many clinicians continue to rely on less empirically supported practices. This disparity highlights a critical gap between research evidence and clinical practice. Research on depression assessments in practice With respect to clinical practice, research has examined how mental health professionals use diagnostic classification systems such as the International Classification of Diseases 37 or the Diagnostic and Statistical Manual of Mental Disorders (DSM) , 38 ,39 what professionals’ attitudes are toward standardized assessment, 40–42 and how clinicians’ perceptions of and approaches to the assessment process are related. 43 , 44 Study Objectives How these methods are applied in the practices of clinicians is less known, and the available evidence suggests a discrepancy between guidelines on standards of practice and those of clinical practice. 8 This study aims to gain insights into the extent to which empirically supported assessments are used by clinicians, with a particular focus on the assessment of depression, given that depression is highly prevalent and the leading cause of disability globally. 45 , 46 Specifically, we investigated (1) collection methods and instruments that clinicians use; (2) the approaches used to integrate and interpret these data (i.e., combination approaches), including clinical and statistical methods; (3) clinicians’ attitudes toward standardized assessment tools and AI-based decision support; and (4) clinicians’ beliefs about the accuracy of their assessments relative to other methods and other clinicians. These questions were explored across diverse clinical settings, professional roles, and countries, with the goal of understanding the extent to which empirically supported assessment practices are implemented in real-world clinical contexts and the factors influencing their adoption. Methods Participants Clinicians were recruited via convenience- and snowball sampling through several different online platforms such as LinkedIN, Facebook groups, emailing lists, workplace platforms (e.g., Slack channels for clinicians), and Prolific (an online platform where participants can be paid to take part). The inclusion criteria were as follows: clinicians who 1) are currently working or have past experience in clinical practice (physically or digitally) and 2) are assessing or treating depression in patients. Clinicians who had never worked with assessing and/or treating depression were excluded ( n = 98). The final sample consisted of 585 participating clinicians: 426 (73%) were female, 153 (26%) were male, and six (1%) indicated ‘other’. There were 162 (28%) clinicians from the United Kingdom, 148 (25%) from the United States, 127 (22%) from Sweden, 105 (18%) from the Netherlands, and 43 (7%) from other countries. Among the different occupations reported, psychologists ( n = 195; 33%) and nurses ( n = 172; 29%) made up the largest part of the sample. Other occupations were physicians ( n = 88; 15%), psychotherapists ( n = 86; 15%), psychiatrists ( n = 22; 4%), and 78 (13%) of the respondents selected ‘other' (see Supplementary Table S1 in the Electronic Supplementary Material [ESM] which presents a cross-tabulation of clinician occupations by country). The clinicians reported a mean of 9.08 years ( SD = 8.2) of experience assessing and/or treating depression. Almost half of the clinicians ( n = 275; 47%) reported assessing and/or treating depression daily or several times a week . The other clinicians reported doing so once a week or less ( n = 118; 20%), once a month or less ( n = 95; 16%), or having done this previously but not anymore ( n = 97; 17%). Material: Survey Content The survey was designed to gather insights into clinicians' practices, attitudes, and experiences regarding depression assessment. It included questions on data collection and combination methods, attitudes toward standardized assessment and AI, self-perceived accuracy, clinical experience, and demographics. The six domains within the questionnaire were (1) Assessment Methods , where clinicians reported on the data collection methods they use (e.g., unstructured, semi-structured, and structured interviews; rating scales), frequency of use, and specific instruments for both initial assessment and follow-up; (2) Data Combination Methods , where clinicians indicated whether they rely on clinical judgment, statistical models, or single data sources to integrate information; (3) Assessment Attitudes , where clinicians were asked about their attitudes regarding the importance of standardized methods for diagnostic assessments, as well as their interest in AI decision support coupled with one of the randomly assigned descriptions (1 = no information on the accuracy of AI, 2 = assuming high accuracy of AI, 3 = empirical data on accuracy of AI); (4) Self-Perceived Accuracy , where clinicians compared their accuracy against peers using similar or different methods, including AI and combinations of methods; (5) Clinical Experience and Confidence , where clinicians reported their experience with depression assessments, their confidence levels, and time spent per assessment; and (6) Demographics , including age, gender, country, profession, therapeutic approach, and employment sector. Detailed descriptions of the questions, response options, and randomized AI descriptions are provided on the ESM, along with the full questionnaire. Procedure Informed consent was obtained from all the clinicians at the start of the survey. They were informed about the nature of the study, that their participation was voluntary and anonymous, that they had the right to withdraw at any time without providing a reason, and that their anonymized answers would be shared according to open science practices. Clinicians completed the survey online, following the question sequence outlined in the Materials section. The clinicians were presented with one of the randomized conditions when asked about their attitudes toward AI decision support. However, owing to an issue with the survey tool, the assigned condition was only recorded for a subset of clinicians, specifically those who completed the survey through Prolific. After completing the survey, the clinicians were debriefed and given the opportunity to send feedback. The survey was available in English, Dutch, or Swedish and took, on average, 10.92 minutes ( SD = 15.89; Med = 7.82) to complete. Statistical Analyses The statistical analyses were performed in R (R Core Team, 2022; for details, see ESM). Descriptive statistics (frequencies [Freq], means [ M ], standard deviations [ SD ], medians [Med]) were calculated for each variable. Pearson correlations ( r ) were used to estimate the relationships between the variables, where the cut-off values for the different strengths were as follows: very weak or no correlation (± .0-.2), weak (± .3-.4), moderate (± .5-.6), strong (± .7-.8), and very strong (± .9 − 1). Analysis of variance (ANOVA) was employed for the AI condition comparisons and the cross-cultural comparisons. All ANOVAs were calculated for unbalanced designs (Type III ANOVA). The assumptions of homogeneity of variances were checked via Levene’s test and were met. For significant main effects, post-hoc pairwise comparisons were conducted with Bonferroni correction to control for multiple testing. The alpha -value was set to 0.05. Results Duration of evaluation The self-reported time it took for clinicians to identify depression in a typical patient ranged from 3 minutes to 15 hours[1] , with a median time of 60 minutes ( SD = 108 minutes, Mean = 87 minutes). Self-reported assessment time did not correlate significantly with the reported use of various data collection methods (see Supplementary Table S2 in the ESM for the correlations between all variables). Aim 1: Data collection methods The clinicians reported which data collection method(s) they used (Fig. 1 a). Most clinicians reported using unstructured interviews (83%) and/or rating scales (83%), followed by semistructured interviews (65%), structured interviews (40%), and/or ‘other’ methods (25%). The clinicians also estimated the proportion of their patients for whom they used these different methods (Fig. 1 b). On average, they reported using clinical interviews for 58% of their patients, followed by rating scales for 51% of their patients, and semistructured interviews for 34% of their patients. Most clinicians (89%) reported using two or more methods across patients (Fig. 1 c). Only 11% of the clinicians reported using a single method across all of their patients, of which the unstructured interviews were the most common method. The most commonly reported data collection method combinations were as follows: 1) Unstructured interviews and rating scales (18%); 2) Unstructured interviews, semistructured interviews, structured interviews, and rating scales (18%); and 3) Unstructured interviews, semistructured interviews, and rating scales ( n = 16%). Most clinicians (84%) indicated performing follow-up assessments; similar patterns regarding the variety of data collection methods were observed among them. A more detailed description is available in the ESM. N = 585 Specific data collection scales and (semi-)structured interviews. Clinicians could indicate which standardized tools and instruments they use. A list of the specific instruments and their corresponding figures is detailed in the ESM and the Mental Health Assessment Dashboard (see MHAD; https://cwiebel.shinyapps.io/ClinicalPractices/ ), an interactive online tool that allows for further analysis and custom comparisons of the data. Aim 2: Data combination methods The vast majority (80%) of the clinicians reported combining the collected data with clinical experience (i.e., the clinical method; Fig. 2 ). Eleven percent reported using the statistical data combination method, and 6% reported not using any data combination method since they used only one source of information. N = 585. Aim 3 : Attitudes towards assessment methods Attitudes toward standardized assessment methods. More than half of the clinicians reported somewhat to strongly agree[ing] (76%) that it is important to use standardized assessment methods with high validity and reliability (Fig. 3 a). A notable number of clinicians reported somewhat- to strongly disagree[ing] (17%) with this statement. Clinicians’ ratings of the importance of using standardized assessments were correlated with their reported frequency of using rating scales ( r = .20, p < .001) and structured interviews ( r = .14, p = .002) but negatively correlated with their reported frequency of using unstructured interviews ( r = − .13, p = .002). Many clinicians have elaborated on their answers. A summary of these elaborations is presented in the SM. N = 585. Attitudes toward AI decision support. More clinicians reported to be probably or definitely interested in receiving AI decision support (43%) than probably not or definitely not (33%; Fig. 3 b). A substantial number reported not knowing or having too little knowledge to take a stand (24%). Greater interest in using AI was correlated with positive agreement on the importance of using standardized assessments ( r = .11, p = .009) and with the reported frequency of using rating scales ( r = .17, p < .001). The attitudes toward the use of AI decision support did not significantly differ between the AI information conditions ( F [2,324] = 1.60, p = .204, n = 359)[2] . Many clinicians have also elaborated on their answers, which are presented on the ESM. Aim 4: Beliefs in one’s confidence and accuracy A majority of the clinicians reported feeling confident (35%), very confident (32%), or extremely confident (10%) in identifying depression, whereas few reported feeling somewhat confident (19%), a little confident (4%), or not at all confident (< 1%). Confidence in identifying depression was positively correlated with the reported assessment frequency of depression ( r = .33, p < .001), work experience with identifying depression ( r = .21, p < .001), and the reported frequency of using rating scales and semistructured interviews as data collection methods ( r = .17, p < .001 and r = .11, p = .007 ). See the ESM for more details. Clinicians reported believing that their assessments (based on the method[s] they usually use for assessing depression) yield superior accuracy compared to those clinicians who use only one data collection method, but not with those clinicians who use a combination of data collection methods (Fig. 4 a). The largest discrepancy between the clinicians’ own estimated assessment accuracy compared to others was noted in comparison to those who relied only on rating scales as their single data collection source. Among the clinicians who reported using more than one method in their assessments (89%), most reported believing their accuracy to be similar to or slightly less than that of other clinicians using multiple methods (Fig. 4 c). However, 80% of clinicians rely on clinical experience to interpret results from multiple sources, and these clinicians reported believing their assessment to be more accurate than that of AI or statistical models (Fig. 4 b). In summary, while clinicians do not consider themselves more accurate than their peers do when a combination of methods is used, they do view their accuracy as superior to that of statistical models. The error bars represent the standard error of the mean. N = 585. Cross-cultural differences Exploratory analyses comparing clinicians from the U.S., the Netherlands, Sweden, and the U.K. identified few small to moderate cross-cultural differences. Clinicians from the U.S. and the U.K. used unstructured interviews for the smallest percentage of patients, with averages of 51% and 46%, respectively. The finding that the vast majority of clinicians report combining collected data with clinical experience is consistent across countries. For a more exploratory cross-cultural analysis, see the ESM and the MHAD . However, these findings should be interpreted with caution due to the heterogeneous nature of the samples, especially with respect to the occupations (e.g., psychologists, psychiatrists, general practitioners, nurses). Discussion This study examined the extent to which mental health professionals use empirically supported methods for assessing depression and their attitudes toward different methods. The results show that depression assessments are conducted by different professions with diverse educational backgrounds, using a wide range of methods and reporting large variations in the time it takes to assess, reflecting real-world practices and the diverse experiences patients encounter. The results show that (1) mental health professionals use a wide variety of data collection methods and that (2) they predominantly rely on the clinical method for data combination. (3) A total of 77% of the clinicians agreed on the importance of using standardized, valid, and reliable assessments, and 43% reported being interested in receiving AI decision support. (4) They are generally confident in their assessments but vary substantially in the reported time it takes to carry out a depression assessment. On average, clinicians report believing they are more accurate than AI or other clinicians are when a single data collection method is used. Interestingly, a wide variety of methods and practices were consistently found across all the examined countries, with only small significant cross-cultural differences. This highlights the paradox of consistent inconsistency in depression assessment practices on an international scale. Data collection methods Clinicians report using different data collection methods with a wide range of combinations and instruments that vary greatly between individuals. Most clinicians (89%) reported using two or more methods, with unstructured interviews and rating scales being the most common. Although a significant portion of the clinicians (16%) did not conduct follow-up assessments, a similar pattern of diverse methods was observed among those who did. This aligns with previous studies highlighting the lack of use of outcome measures. 47 , 48 While unstructured interviews have shown low interrater reliability for diagnosing depression and are seldom used in research, 29,49 various rating scales for assessing depression differ significantly (i.e., they capture different symptoms of depression and only partly overlap 30 , 50 ; and are often used interchangeably in research. 30 , 49 – 51 This variation in data collection methods may be problematic for providing equal, standardized, high-quality care, especially since research shows that different methods and instruments yield varying results. The absence of a single industry standard for diagnostic manuals further contributes to inconsistencies in assessments and diagnoses. 51 Data combination methods While approximately one in ten clinicians reported using the statistical method to form their final assessment, 80% reported interpreting and weighing the collected data on the basis of their clinical experience and judgment. Despite extensive research showing that the statistical method for combining information is as good as or superior to the clinical judgment of an expert clinician, 6,13,24,36 the majority of final assessments still rely on the clinical method. As early as 35 years ago, Dawes et al. 14 argued for the routine use of statistical methods in clinical assessment and decision-making. Grove and Meehl 21 addressed this issue, concluding that using the less accurate of two assessment methods “is not only unscientific and irrational, it is unethical” ( p . 323). However, possible explanations for the continued reliance on clinical data combination might be the lack of easily implemented, user-friendly, and validated statistical models that comply with the required regulations (e.g., CE-marking and FDA approval). Clinical validation and readiness for implementation should be weighed heavily before employing statistical models since, for example, flawed algorithms trained on small or biased training data can have adverse effects. 52 This concern is further amplified by widespread apprehensions about the safety and validity of AI models, as highlighted by clinicians in this survey. Attitudes toward standardized assessment A sizable group of clinicians, approximately one-fifth, report that they somewhat - to strongly disagree with the importance of using standardized assessment methods that have been demonstrated in research to have high validity and reliability — which we expected to find a unanimous agreement on. Even though 76% of the clinicians reported that they somewhat- to strongly agree with the statement, the importance they assigned to standardized assessment methods did not affect the final decision for most clinicians, who still use a combination of clinical data methods. This corresponds to other research examining clinicians’ attitudes, which have shown a mostly, but not completely, positive attitude towards standardized assessment methods: Clinicians seem to be positive about the quality of standardized assessment methods but less positive about their usefulness in practice, and the advantages of standardized assessment over clinical judgment. 40 – 42 To enable better implementation, the need to overcome specific barriers and issues, such as adequate training and practicality, is highlighted. 40 – 42 , 44 Attitudes toward AI Clinicians’ interest in AI decision support for depression assessment is mixed: 43% express interest, 33% show disinterest, and 23% remain undecided due to insufficient information. This highlights the need to update clinicians on AI advancements and benefits and involve them in algorithm development and validation. Key concerns, including accountability, transparency, and privacy, must be addressed. 53 Interest in AI support correlates with the use of rating scales and positive attitudes toward standardized methods, suggesting that clinicians familiar with these practices could facilitate AI integration in clinical settings. However, enhanced education on AI and statistical models is essential. For example, a study on psychiatrists’ views of generative AI found widespread use for administrative tasks but skepticism toward its active role in patient care. 54 The lack of differences in attitudes across AI descriptions and diverse responses in open comments further emphasize the need for comprehensive educational initiatives to foster an understanding of AI’s capabilities, limitations, and ethical integration into mental health practices. Clinicians Assessment Accuracy Strikingly, a majority of clinicians who rely on a combination of clinical data believe that their assessments are more accurate than those derived from a statistical method or AI, which contradicts findings from previous research. 6 , 13 , 24 , 36 The clinicians ranked the accuracy of their assessments as superior to those conducted by peers who use only a single method (particularly rating scales alone) but not compared with those using a combination of methods. These findings suggest a general belief among clinicians in the superiority of employing multiple methods for data collection and combining data via a clinical method. The belief in the clinical method could be explained by a more general tendency of individuals to overestimate their cognitive abilities. 55 However, clinicians’ confidence and accuracy beliefs have been shown to be poor indicators of their assessment accuracy, 20,23,56 which is troubling given the impact it may have on patients and their course of treatment. Limitations The results are constrained by certain limitations, yet they indicate several directions for future research to further our understanding of depression assessment practices and facilitate the adoption of empirically supported assessment methods. Given the study’s broad focus on depression assessments in real-world clinical settings, the sample comprises a wide range of occupations, including psychologists, psychotherapists, psychiatrists, and general practitioners. However, this makes it challenging to understand the experiences of specific occupations; therefore, future research could focus on single professional groups to gain deeper insights into their unique practices and challenges. Similarly, the goals and methods of clinicians’ assessments can differ significantly both between and within countries, occupations, and clinics, all of which warrant further investigation. This study included only Western countries, each with distinct guidelines, rules, and regulations governing mental health practices. The distribution of occupations also varies across countries, which necessitates caution when interpreting cross-country differences. Future research should expand to include non-Western countries to gain insights into how mental health practices are carried out in diverse cultural and regulatory contexts. Additionally, this study specifically examined the assessment of depression, so the findings may not be generalizable to other mental health conditions. The use of a self-report survey format introduces the potential for recall bias and may reflect clinicians’ perceptions of best practices rather than their actual behaviors. As a result, the data could overestimate the extent to which empirically supported assessment methods are employed in practice. Finally, online recruitment and sampling may have restricted participation to specific subgroups. The sampling strategy is also likely to favor digitally active clinicians, which should be considered when interpreting the findings. In addition, AI-related randomization data were missing for some participants, which reduced the statistical power. Future research could benefit from random sampling directly from clinical settings or from observational studies of audits of actual practices, etc. Potential implications and future research A better understanding of current depression assessment practices can provide clinicians, policymakers, and researchers with a foundation for improving healthcare quality. Understanding current practices can support clinicians, regulatory bodies, and researchers in joining forces to lead changes and update guidelines and policies. The Mental Health Assessment Dashboard can inform discussions on the current ( limited ) use of empirically supported assessments and any future integration of AI in clinical settings. Furthermore, there is a need for more research regarding the potential risks of current depression assessment practices and their implications for patients and the challenges for clinicians in embracing newer empirically supported methods. 15 , 57 There is a lack of research on how clinicians combine assessment methods and the potential shortcomings associated with the clinical combination method. The variability in assessors and assessment methods reflects real-world practice and highlights the need for clearer guidelines and training to ensure equitable and consistent standardization. In other words, the outcome of a depression assessment for depression should not depend on where, when, or by whom the patient is assessed. Additionally, many clinicians’ concerns about AI should be addressed, including biases and ethical issues, as well as how they can work with AI. Transparent communication and education about the capabilities and limitations of AI, along with evidence from robust studies, can potentially build trust and acceptance among clinicians. 58 Conclusion Previous research 7 , 51 has typically studied depression assessment methods in research settings or focused on attitudes in clinical settings 40 – 42 rather than actual assessment practices in different clinical practices. Our results highlight heterogeneity in depression assessments: Many different professionals conduct depression assessments in clinical practice (e.g., psychiatrists, psychologists, general practitioners, nurses); employ a wide variety of methods (ranging from unstructured to structured approaches using numerous tools) in different combinations, and report a great difference in the time it takes to carry out the assessment. The results reveal a pattern of skepticism toward AI, underestimation of statistical models, and overestimation of one’s own accuracy, which aligns with findings from previous studies. While clinicians express confidence in their ability to assess depression, the widespread use of unstructured interviews alongside standardized tools, coupled with a predominant reliance on clinical judgment over statistical models, suggests a disconnect between best practices supported by research and everyday clinical practice. Adopting empirically supported practices, including standardized, valid, and reliable data collection and combination assessment methods, is essential to synthesize information systematically and improve the quality of the data used in clinical judgment. This alignment between clinical practice and research evidence would enhance patient care while allowing clinicians to consider individual factors. However, achieving this requires harmonizing clinical judgment with data-driven approaches, as well as aligning manuals and guidelines with empirical research. Future initiatives should focus on developing accessible, practical, validated tools and training programs that bridge the gap between research and practice, fostering integration and consistency in depression assessment. Abbreviations AI Artificial Intelligence DSM Diagnostic and Statistical Manual of Mental Disorders ESM Electronic Supplementary Material ICD International Classification of Diseases LLM Large Language Model (only include if mentioned earlier; I didn’t sit explicitly) MHAD Mental Health Assessment Dashboard U.K. United Kingdom U.S. United States WHO World Health Organization Declarations Ethics approval and consent to participate The study recieved ethical approval from Leiden University (2022-08-10-M.L. Molendijk-V5-4037) and was exempt from ethical approval at the University of Pennsylvania (United States). The study is deemed exempt from requiring ethical approval according to Swedish Law (s§3-4 of the Act [2003:460] on the ethical review of research involving humans in Sweden). Consent for publication Not applicable Availability of data and materials The datasets generated and/or analysed during the current study are available at https://osf.io/pw8mk/. The code for the Mental Health Assessment Dashboard is also shared here in case the app no longer can be hosted (https://cwiebel.shinyapps.io/ClinicalPractices/). Competing interests O.N.E. Kjell and K. Kjell co-founded a start-up that uses computational language assessments to assess mental health problems. E.C. Stade has received fees for advising a start-up that uses computational language assessment to measure psychological constructs. The other authors report having no conflicts of interest with respect to the contents, authorship, or publication of this article. Funding O.N.E. Kjell, K. Kjell and V.C. Eijsbroek were funded by FORTE (2022-01022). Authors' contributions Conceptualization: Ideas; formulation or evolution of overarching research goals and aims. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell Data curation: Management activities to annotate (produce metadata), scrub data and maintain research data (including software code, where it is necessary for interpreting the data itself) for initial use and later re-use. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel Formal analysis: Application of statistical, mathematical, computational, or other formal techniques to analyze or synthesize study data. Clara Wiebel, Veerle C. Eijsbroek, Damon Navandi, and Oscar N. E. Kjell Funding acquisition: Acquisition of the financial support for the project leading to this publication. Katarina Kjell, Oscar N. E. Kjell Investigation: Conducting a research and investigation process, specifically performing the experiments, or data/evidence collection. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell Methodology: Development or design of methodology; creation of models. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell Project administration: Management and coordination responsibility for the research activity planning and execution. Katarina Kjell, Oscar N. E. Kjell Resources: Provision of study materials, reagents, materials, patients, laboratory samples, animals, instrumentation, computing resources, or other analysis tools. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell Software: Programming, software development; designing computer programs; implementation of the computer code and supporting algorithms; testing of existing code components. Clara Wiebel Supervision: Oversight and leadership responsibility for the research activity planning and execution, including mentorship external to the core team. Oscar N. E. Kjell Validation: Verification, whether as a part of the activity or separate, of the overall replication/reproducibility of results/experiments and other research outputs. Visualization: Preparation, creation and/or presentation of the published work, specifically visualization/data presentation. Clara Wiebel, Veerle C. Eijsbroek, Damon Navandi, Oscar N. E. Kjell Writing—original draft: Preparation, creation and/or presentation of the published work, specifically writing the initial draft (including substantive translation). Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Oscar N. E. Kjell Writing—review and editing: Preparation, creation and/or presentation of the published work by those from the original research group, specifically critical review, commentary or revision: including pre- or post-publication stages. Damon Navandi, Veerle C. Eijsbroek, Clara Wiebel Acknowledgements Not applicable References Shen H, Zhang L, Xu C, Zhu J, Chen M, Fang Y. Analysis of Misdiagnosis of Bipolar Disorder in An Outpatient Setting. Shanghai Archives of Psychiatry [Internet]. 2018;30(2):93–101. Available from: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5936046/ Stahnke B. A systematic review of misdiagnosis in those with obsessive-compulsive disorder. Journal of Affective Disorders Reports [Internet]. 2021;6(1):100231. Available from: https://www.sciencedirect.com/science/article/pii/S2666915321001578 Fried EI, Proppert RKK, Rieble CL. Building an Early Warning System for Depression: Rationale, Objectives, and Methods of the WARN-D Study. Clinical Psychology in Europe [Internet]. 2023;5(3):1–25. Available from: https://cpe.psychopen.eu/index.php/cpe/article/view/10075 Eijsbroek VC, Katarina Kjell, Schwartz A, Boehnke JR, Fried EI, Klein DN et al. The LEADING Guideline. Reporting Standards for Expert Panel, Best-Estimate Diagnosis, and Longitudinal Expert All Data (LEAD) Studies. medRxiv (Cold Spring Harbor Laboratory). 2024. Hunsley J, Mash EJ. Evidence-based assessment. Annual review of clinical psychology [Internet]. 2007;3:29–51. Available from: https://www.ncbi.nlm.nih.gov/pubmed/17716047 Meehl PE. Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. Minneapolis: University of Minnesota Press; 1954. Pettersson A, Boström KB, Gustavsson P, Ekselius L. Which instruments to support diagnosis of depression have sufficient accuracy? A systematic review. Nord J Psychiatry. 2015;69(7):497–508. Cook JR, Hausman EM, Jensen-Doss A, Hawley KM. Assessment Practices of Child Clinicians. Assessment. 2016;24(2):210–21. Graham S, Depp C, Lee EE, Nebeker C, Tu X, Kim HC, et al. Artificial Intelligence for Mental Health and Mental Illnesses: an Overview. Curr Psychiatry Rep. 2019;21(11):116. Oscar NE, Kjell K, Kjell H. Andrew Schwartz. Beyond Rating Scales: With Targeted Evaluation, Language Models are Poised for Psychological Assessment. Psychiatry Res. 2023;333:115667–7. Lee EE, Torous J, De Choudhury M, Depp CA, Graham SA, Kim HC, et al. Artificial Intelligence for Mental Healthcare: Clinical Applications, Barriers, Facilitators, and Artificial Wisdom. Biol Psychiatry: Cogn Neurosci Neuroimaging. 2021;6(9):856–64. Mueller AE, Segal DL. Structured versus Semistructured versus Unstructured Interviews. Encyclopedia Clin Psychol. 2015;1(7):1–7. Ægisdóttir S, White MJ, Spengler PM, Maugherman AS, Anderson LA, Cook RS, et al. The Meta-Analysis of Clinical Judgment Project: Fifty-Six Years of Accumulated Research on Clinical Versus Statistical Prediction. Couns Psychol. 2006;34(3):341–82. Dawes R, Faust D, Meehl P. Clinical versus actuarial judgment. Science [Internet]. 1989;243(4899):1668–74. Available from: http://meehl.umn.edu/sites/meehl.dl.umn.edu/files/138cstixdawesfaustmeehl.pdf DeRubeis RJ. The history, current status, and possible future of precision mental health. Behav Res Ther. 2019;123:103506. Shatte ABR, Hutchinson DM, Teague SJ. Machine learning in mental health: a scoping review of methods and applications. Psychological Medicine [Internet]. 2019;49(09):1426–48. Available from: https://www.cambridge.org/core/journals/psychological-medicine/article/abs/machine-learning-in-mental-health-a-scoping-review-of-methods-and-applications/0B70B1C827B3A4604C1C01026049F7D9 Bennett CC, Hauser K. Artificial intelligence framework for simulating clinical decision-making: A Markov decision process approach. Artif Intell Med. 2013;57(1):9–19. Egon Brunswik, Cooksey RW. The Conceptual Framework of Psychology [1952]. 2001;225–37. Hammond KR, Hursch CJ, Todd FJ. Analyzing the components of clinical inference. Psychol Rev. 1964;71(6):438–56. Karelaia N, Hogarth RM. Determinants of linear judgment: A meta-analysis of lens model studies. Psychol Bull. 2008;134(3):404–26. Grove WM, Meehl PE. Comparative efficiency of informal (subjective, impressionistic) and formal (mechanical, algorithmic) prediction procedures: The clinical-statistical controversy. Psychol Public Policy Law. 1996;2(2):293–323. Bowes SM, Ammirati RJ, Costello TH, Basterfield C, Lilienfeld SO. Cognitive biases, heuristics, and logical fallacies in clinical practice: A brief field guide for practicing clinicians and supervisors. Prof Psychology: Res Pract. 2020;51(5):435–45. Saposnik G, Redelmeier D, Ruff CC, Tobler PN. Cognitive biases associated with medical decisions: a systematic review. BMC Med Inf Decis Mak. 2016;16(1). Garb HN, Wood JM. Methodological advances in statistical prediction. Psychol Assess. 2019;31(12):1456–66. Collins FS, Varmus H. A New Initiative on Precision Medicine. N Engl J Med. 2015;372(9):793–5. Davenport T, Kalakota R. The Potential for Artificial Intelligence in Healthcare. Future Healthcare Journal [Internet]. 2019;6(2):94–8. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC6616181/ Ali GC, Ryan G, De Silva MJ. Validated Screening Tools for Common Mental Disorders in Low and Middle Income Countries: A Systematic Review. Burns JK, editor. PLOS ONE. 2016;11(6):e0156939. Breedvelt JJF, Zamperoni V, South E, Uphoff EP, Gilbody S, Bockting CLH, et al. A systematic review of mental health measurement scales for evaluating the effects of mental health prevention interventions. Eur J Pub Health. 2020;30(3):510–6. Miller PR, Dasher R, Collins R, Griffiths P, Brown F. Inpatient diagnostic assessments: 1. Accuracy of structured vs. unstructured interviews. Psychiatry Res. 2001;105(3):255–64. Fried EI. The 52 symptoms of major depression: Lack of content overlap among seven common depression scales. J Affect Disord. 2017;208(208):191–7. Chevance A, Ravaud P, Tomlinson A, Le Berre C, Teufer B, Touboul S, et al. Identifying outcomes for depression that matter to patients, informal caregivers, and health-care professionals: qualitative content analysis of a large international online survey. Lancet Psychiatry. 2020;7(8):692–702. Zimmerman M, Martinez JH, Friedman M, Boerescu DA, Attiullah N, Toba C. How Can We Use Depression Severity to Guide Treatment Selection When Measures of Depression Categorize Patients Differently? J Clin Psychiatry. 2012;73(10):1287–91. Shafer AB. Meta-analysis of the factor structures of four depression questionnaires: Beck, CES-D, Hamilton, and Zung. J Clin Psychol. 2005;62(1):123–46. van Loo HM, de Jonge P, Romeijn JW, Kessler RC, Schoevers RA. Data-driven subtypes of major depressive disorder: a systematic review. BMC Med. 2012;10(1). Fried EI, van Borkulo CD, Epskamp S, Schoevers RA, Tuerlinckx F, Borsboom D. Measuring depression over time.. Or not? Lack of unidimensionality and longitudinal measurement invariance in four common rating scales of depression. Psychol Assess. 2016;28(11):1354–67. Grove WM, Zald DH, Lebow BS, Snitz BE, Nelson C. Clinical versus mechanical prediction: A meta-analysis. Psychol Assess. 2000;12(1):19–30. ICD-11 [Internet]. icd.who.int. Available from: https://icd.who.int American Psychiatric Association. Diagnostic and statistical manual of mental disorders. Diagnostic and Statistical Manual of Mental Disorders. 5th ed. 2013;5(5). First MB, Rebello TJ, Keeley JW, Bhargava R, Dai Y, Kulygina M, et al. Do mental health professionals use diagnostic classifications the way we think they do? A global survey. World Psychiatry. 2018;17(2):187–95. Danielson M, Månsdotter A, Fransson E, Dalsgaard S, Larsson J-O. Clinicians’ attitudes toward standardized assessment and diagnosis within child and adolescent psychiatry. Child Adolesc Psychiatry Mental Health. 2019;13(1). Jensen-Doss A, Hawley KM. Understanding Barriers to Evidence-Based Assessment: Clinician Attitudes Toward Standardized Assessment Tools. Journal of Clinical Child & Adolescent Psychology [Internet]. 2010;39(6):885–96. Available from: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3058768/ Jensen-Doss A, Hawley KM. Understanding Clinicians’ Diagnostic Practices: Attitudes Toward the Utility of Diagnosis and Standardized Diagnostic Tools. Administration and Policy in Mental Health and Mental Health Services Research [Internet]. 2011;38(6):476–85. Available from: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6114089/ Bohman B. Clinicians’ perceptions and practices of diagnostic assessment in psychiatric services. BMC Psychiatry. 2023;23(1). Cho E, Tugendrajch SK, Marriott BR, Hawley KM. Evidence-Based Assessment in Routine Mental Health Services for Youths. Psychiatric Serv. 2020;72(3): appi.ps.2019005 Mathers CD, Loncar D. Projections of Global Mortality and Burden of Disease from 2002 to 2030. Samet J, editor. PLoS Medicine [Internet]. 2006;3(11):e442. Available from: https://pubmed.ncbi.nlm.nih.gov/17132052/ Prevalence. and variability of current depressive disorder in 27 European countries: a population-based study. The Lancet Public Health [Internet]. 2021; Available from: https://www.sciencedirect.com/science/article/pii/S2468266721000475#bib9 Boswell JF, Kraus DR, Miller SD, Lambert MJ. Implementing routine outcome monitoring in clinical practice: Benefits, challenges, and solutions. Psychother Res. 2013;25(1):6–19. Gilbody SM, House AO, Sheldon TA. Psychiatrists in the UK do not use outcomes measures. Br J Psychiatry. 2002;180(2):101–3. Regier DA, Narrow WE, Clarke DE, Kraemer HC, Kuramoto SJ, Kuhl EA, et al. DSM-5 Field Trials in the United States and Canada, Part II: Test-Retest Reliability of Selected Categorical Diagnoses. Am J Psychiatry. 2013;170(1):59–70. Newson JJ, Hunter D, Thiagarajan TC. The heterogeneity of mental health assessment. Frontiers in Psychiatry [Internet]. 2020;11(76):1–24. Available from: https://www.frontiersin.org/articles/ 10.3389/fpsyt.2020.00076/full Santor DA, Gregus M, Welch A. FOCUS ARTICLE: Eight Decades of Measurement in Depression. Measurement: Interdisciplinary Research & Perspective. 2006;4(3):135–55. Topol EJ. High-performance medicine: the Convergence of Human and Artificial Intelligence. Nature Medicine [Internet]. 2019;25(1):44–56. Available from: https://www.nature.com/articles/s41591-018-0300-7 Leslie D, systems in the public sector Dr David Leslie Public Policy Programme. Understanding artificial intelligence ethics and safety A guide for the responsible design and implementation of AI. Understanding artificial intelligence ethics and safety [Internet]. 2019; Available from: https://www.turing.ac.uk/sites/default/files/2019-06/understanding_artificial_intelligence_ethics_and_safety.pdf Blease C, Worthen A, Torous J. Psychiatrists’ Experiences and Opinions of Generative Artificial Intelligence in Mental Healthcare: An Online Mixed Methods Survey. Psychiatry Res. 2024;333:115724–4. Heck PR, Simons DJ, Chabris CF. 65% of Americans believe they are above average in intelligence: Results of two nationally representative surveys. van Amelsvoort T, editor. PLOS ONE. 2018;13(7):e0200103. Miller DJ, Spengler ES, Spengler PM. A meta-analysis of confidence and judgment accuracy in clinical decision making. J Couns Psychol. 2015;62(4):553–67. Betancourt TS, Chambers DA. Optimizing an Era of Global Mental Health Implementation Science. JAMA Psychiatry. 2016;73(2):99. Misra R, Keane PA, Hogg HDJ. How should we train clinicians for artificial intelligence in healthcare? Future Healthc J. 2024;11(3):100162. Footnotes Nine cases with a reported time greater than the mean plus 3 standard deviations were seen as outliers and thus removed from the time analysis. Their reported assessment times ranged from 20 to 60 hours, respectively. Additionally, seven cases reported 0 minutes; these were removed because it is possible that they do not assess depression at all. The lower number of clinicians in this analysis is due to the randomisation condition not being saved for all clinicians. Additional Declarations Competing interest reported. O.N.E. Kjell and K. Kjell co-founded a start-up that uses computational language assessments to assess mental health problems. E.C. Stade has received fees for advising a start-up that uses computational language assessment to measure psychological constructs. The other authors report having no conflicts of interest with respect to the contents, authorship, or publication of this article. Supplementary Files SupplementaryMaterial.docx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 23 Mar, 2026 Reviews received at journal 19 Mar, 2026 Reviewers agreed at journal 19 Mar, 2026 Reviews received at journal 09 Mar, 2026 Reviewers agreed at journal 01 Mar, 2026 Reviewers agreed at journal 27 Feb, 2026 Reviewers invited by journal 27 Feb, 2026 Editor invited by journal 26 Feb, 2026 Editor assigned by journal 25 Feb, 2026 Submission checks completed at journal 25 Feb, 2026 First submitted to journal 23 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8951240","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":600221311,"identity":"583036ec-ab89-4349-9a76-a77c323b80b7","order_by":0,"name":"Damon Navandi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7ElEQVRIiWNgGAWjYDCCA0hsiYQKIMnM3ECKljMgLYykaGFsA1EEtPAdP534uIDhjrzB8bMHbzycVxvN3w7U8qNiG04tkmdyNxvPYHhmuOFMXrJF4rbjuTMOMzYw9py5jVOLwYHcbdI8DIcZZzbkmEkkbjuW2wDUwszYhkfL+bdgLfYz+98Atcw5ljufoJYbEFsS+yVAtjTU5G4gpEXyxtvNxjwGh5P7Jd4YWyQcO5C7EajlID6/8J3P3fiYp+KwbRt/juHNHzV1ufPOHz744EcFbi1Q58FZh8HkAQLqUUAdKYpHwSgYBaNghAAA+9NfAgVBME8AAAAASUVORK5CYII=","orcid":"","institution":"Lund University","correspondingAuthor":true,"prefix":"","firstName":"Damon","middleName":"","lastName":"Navandi","suffix":""},{"id":600221312,"identity":"fe0071b9-272e-4d10-92b5-2f8f498965b0","order_by":1,"name":"Veerle C. Eijsbroek","email":"","orcid":"","institution":"Lund University","correspondingAuthor":false,"prefix":"","firstName":"Veerle","middleName":"C.","lastName":"Eijsbroek","suffix":""},{"id":600221313,"identity":"f0215419-2b15-437d-aadf-265b33ea1720","order_by":2,"name":"Clara Wiebel","email":"","orcid":"","institution":"Lund University","correspondingAuthor":false,"prefix":"","firstName":"Clara","middleName":"","lastName":"Wiebel","suffix":""},{"id":600221314,"identity":"300a183b-ef35-4fac-a306-370a4cfb830f","order_by":3,"name":"Katarina Kjell","email":"","orcid":"","institution":"Lund University","correspondingAuthor":false,"prefix":"","firstName":"Katarina","middleName":"","lastName":"Kjell","suffix":""},{"id":600221315,"identity":"af8976ae-77a8-4be3-8366-cc05ede9cba1","order_by":4,"name":"Marc L. Molendijk","email":"","orcid":"","institution":"Leiden University","correspondingAuthor":false,"prefix":"","firstName":"Marc","middleName":"L.","lastName":"Molendijk","suffix":""},{"id":600221316,"identity":"4eb51a0c-0f7f-42c3-8a1f-8f1dcbc708de","order_by":5,"name":"Elizabeth C. Stade","email":"","orcid":"","institution":"Stanford University","correspondingAuthor":false,"prefix":"","firstName":"Elizabeth","middleName":"C.","lastName":"Stade","suffix":""},{"id":600221317,"identity":"46aacd8a-f824-4bf5-9387-bd10486906a7","order_by":6,"name":"Oscar N. E. Kjell","email":"","orcid":"","institution":"Lund University","correspondingAuthor":false,"prefix":"","firstName":"Oscar","middleName":"N. E.","lastName":"Kjell","suffix":""}],"badges":[],"createdAt":"2026-02-24 00:08:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8951240/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8951240/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104401262,"identity":"5d9ba044-9899-4c08-bf23-2e18019b5fa3","added_by":"auto","created_at":"2026-03-11 12:12:14","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":270295,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of depression assessment methods, including data collection and combination methods.\u003cbr\u003e\n \u003c/strong\u003e\u003cem\u003eNote. \u003c/em\u003eThe data combination method does not necessarily require data from different data collection methods; for example, responses to items of a single rating scale can be aggregated and interpreted via a statistical model (following an aggregation schema or using an established cut-off point) or interpreted via a clinical method in which individual items are examined via clinical experience. In the statistical method, it is also feasible to include data from a clinical (unstructured) interview in a statistical model (e.g., to what extent the patient responds, their response times, and behaviors).\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/da6f9f7f5923396df9efcd9c.png"},{"id":104401616,"identity":"0a68bc34-6339-42d6-8de8-a5ac053ca43c","added_by":"auto","created_at":"2026-03-11 12:13:09","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":304967,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 1 A‒C | Clinicians’ reported data collection methods: A) number (percentage) of clinicians using each method, B) mean percentage of patients each method is used for and C) the number of methods used per clinician across all patients.\u003cbr\u003e\n \u003c/strong\u003e\u003cem\u003eN\u003c/em\u003e = 585\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/6f5405a8bda5be643d224d07.png"},{"id":103896300,"identity":"9bca69fa-1f1f-4d5a-86d9-0a5e86bd82b2","added_by":"auto","created_at":"2026-03-04 09:00:13","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":94445,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 2 | Clinicians’ reported use of different data combination methods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eN \u003c/em\u003e= 585.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/2efb5a57648819025def1b95.png"},{"id":103896304,"identity":"f256aaf4-c2d5-49e5-8bdc-b25dfba393aa","added_by":"auto","created_at":"2026-03-04 09:00:13","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":183512,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 3 A-B | Clinicians' reported attitudes toward the use of A) standardized assessment methods and B) AI decision support.\u003c/strong\u003e \u003cbr\u003e\n \u003cem\u003eN\u003c/em\u003e = 585.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/2b81d7edce2ef179782932a6.png"},{"id":104401213,"identity":"03b5a16d-f788-46c7-8b58-8cfa783361e3","added_by":"auto","created_at":"2026-03-11 12:12:07","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":158730,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 4 A‒C | Clinicians’ self-perceived assessment accuracy: A) compared with that of clinicians using other methods, B) compared with that of clinicians using AI, and C) compared with that of clinicians using a combination of methods.\u003c/strong\u003e\u003cbr\u003e\nThe error bars represent the standard error of the mean. \u003cem\u003eN\u003c/em\u003e = 585.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/92fd5cf17b4c37681eeff9d5.png"},{"id":104408373,"identity":"e99c3d9d-6d89-4567-997f-fba01b1c04db","added_by":"auto","created_at":"2026-03-11 12:42:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2469200,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/07d3a1c3-1672-42dc-ac3c-9a5fa724c614.pdf"},{"id":103896302,"identity":"d1e4aeec-6a14-40f6-abd2-2c2a0f0e1226","added_by":"auto","created_at":"2026-03-04 09:00:13","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":975846,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-8951240/v1/d30f2ee5370b5f55eb368726.docx"}],"financialInterests":"Competing interest reported. O.N.E. Kjell and K. Kjell co-founded a start-up that uses computational language assessments to assess mental health problems. E.C. Stade has received fees for advising a start-up that uses computational language assessment to measure psychological constructs. The other authors report having no conflicts of interest with respect to the contents, authorship, or publication of this article.","formattedTitle":"Clinicians’ Practices and Attitudes Toward Depression Assessment: A Cross-Sectional Survey Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAccurate assessment of mental health disorders is crucial for early detection and effective treatment planning. Inaccurate assessment of the type and severity of mental disorders can have significant negative impacts on people\u0026rsquo;s lives, highlighting the importance of precise assessment methods for effective prevention, early detection, and appropriate treatment.\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e There are many different methods used to assess mental health problems, including various \u003cem\u003edata collection methods\u003c/em\u003e such as interviews and rating scales, as well as various \u003cem\u003edata combination methods\u003c/em\u003e, i.e., approaches for evaluating, integrating, and interpreting different sources of information (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThe validity and reliability of depression assessment methods have been extensively studied,\u003csup\u003e6,7\u003c/sup\u003e with some methods being more reliable and valid than others are. However, less is known about the practices of mental health professionals, and the available evidence suggests a discrepancy between guidelines on standards of practice and those on clinical practice.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e This study aims to gain insights into the extent to which empirically supported assessments are used by mental health professionals, as well as their attitudes toward these methods, which is critical for identifying risks and challenges in current assessment practices and the implementation of new assessment methods. Our study may also add to the discourse on integrating artificial intelligence-based (AI) decision-support systems into assessment processes.\u003csup\u003e\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eNote\u003c/strong\u003e \u003cp\u003eThe data combination method does not necessarily require data from different data collection methods; for example, responses to items of a single rating scale can be aggregated and interpreted via a statistical model (following an aggregation schema or using an established cut-off point) or interpreted via a clinical method in which individual items are examined via clinical experience. In the statistical method, it is also feasible to include data from a clinical (unstructured) interview in a statistical model (e.g., to what extent the patient responds, their response times, and behaviors).\u003c/p\u003e \u003c/p\u003e"},{"header":"Data collection methods","content":"\u003cp\u003eData collection methods range from \u003cem\u003eunstructured\u003c/em\u003e to \u003cem\u003estructured\u003c/em\u003e.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e \u003cem\u003eUnstructured data collection methods\u003c/em\u003e include \u003cem\u003eunstructured clinical interviews\u003c/em\u003e, which allow clinicians to explore patients\u0026rsquo; issues freely and adjust the direction of the assessment on the basis of patient\u0026rsquo;s responses and behaviors without predetermined questions. \u003cem\u003eStructured data collection methods\u003c/em\u003e are highly organized and follow a predetermined format, ensuring that the same questions are asked of each person. Structured methods aim to yield more reliable and comparable results and are especially well suited for assessing a set of validated criteria for a disorder or clinical phenomenon (e.g., evaluating whether a patient meets a diagnostic manual's criteria for a major depressive episode). Examples include rating scales and structured clinical interviews, as well as semistructured clinical interviews where questions are predetermined, but the interviewer is allowed to ask additional questions or exclude ones to adjust the sourcing of information on a case-to-case basis. The interviews were conducted by clinicians, while for the rating scales, some were clinician-rated, and others were self-reported.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData combination methods\u003c/h2\u003e \u003cp\u003eTwo primary approaches for evaluating, integrating, and interpreting different sources of information have been identified for depression assessment: the \u003cem\u003eclinical\u003c/em\u003e approach and the \u003cem\u003estatistical\u003c/em\u003e approach.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e \u003cem\u003eThe clinical data combination method\u003c/em\u003e is based on human judgment. Clinicians assess various sources of information on the basis of personal judgment, clinical experience, theoretical perspectives as well as surrounding and individual factors (e.g., subtle deviations from a patient\u0026rsquo;s baseline behavior, appearance, overt signs such as direct statements of distress or crying, or general deduction on the basis of the overall information gathered).\u003c/p\u003e \u003cp\u003eIn contrast, \u003cem\u003ethe statistical data combination method\u003c/em\u003e makes use of a statistical model that is based on empirically established relationships between the different sources of information and the outcome of interest. This can, for example, involve the clinician entering client data into formulas, actuarial tables, or charts.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e To establish relationships among variables, statistical models may be constructed via relatively simple statistical approaches such as linear or logistic regression. It can also involve more complex statistical techniques, such as AI, which includes techniques such as machine learning and natural language processing. Statistical data combination methods can, for example, be used for screening, comprehensive diagnostic assessment support, treatment personalization, or symptom monitoring.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eA growing body of research supports the use of AI in mental health care, demonstrating its ability to improve the accuracy and efficiency of diagnostic assessments.\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e Given its promising applications and increasing capabilities, we focus on AI as a key component of the statistical approach in the following sections.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFinal assessment and decision-making\u003c/h3\u003e\n\u003cp\u003eIn the data combination phase, both the clinical and the statistical data combination methods can be used at different stages (see Brunswick Lens Model).\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e For example, a clinician may use a statistical method to integrate a patient\u0026rsquo;s rating scale responses but rely on a clinical method to combine those results with information from an unstructured interview. However, when the results from the two different methods disagree, one cannot follow both of them: When they \u003cem\u003edisagree\u003c/em\u003e, the final decision will necessarily have to rely on one method over the other.\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eNavigating and integrating different information via the clinical data combination method can be cognitively strenuous and complex,\u003csup\u003e22,23\u003c/sup\u003e especially when the various sources are in conflict with one another. The clinical method has been shown to be less reliable than statistical models.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e However, the implementation of statistical models can be difficult due to insufficient guidelines, education, or resources, such as a lack of validated advanced statistical models available for practical use in psychological assessments.\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e For example, models that enable the statistical combination of data from several rating scales and clinical history are still uncommon.\u003c/p\u003e \u003cp\u003eSince the introduction of the clinical versus statistical controversy in the 1950s, statistical modeling has undergone significant improvements as a result of advancements in the field of AI.\u003csup\u003e9,10,24\u003c/sup\u003e Statistical models, often including AI technology, have become increasingly prevalent in healthcare, especially in medicine.\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e Nevertheless, the question remains whether the benefits of such instruments are used for depression assessments in clinical practice and, if so, to what extent, in what way, and how clinicians\u0026rsquo; attitudes reflect their tendencies to use the statistical data combination methods.\u003c/p\u003e\n\u003ch3\u003eResearch support for assessment methods\u003c/h3\u003e\n\u003cp\u003eThe validity and reliability of depression assessment methods have been systematically researched for many decades. Semistructured or structured interviews are often regarded as the benchmark standard for depression assessment because they demonstrate greater validity and reliability than unstructured interviews do.\u003csup\u003e7,27,28\u003c/sup\u003e In contrast, unstructured clinical interviews, although widely used in practice, are less reliable and less valid,\u003csup\u003e12,29\u003c/sup\u003e and they are rarely used in research. Self-report questionnaires, while highly reliable, have other shortcomings. For example, rating scales targeting depression: \u003cb\u003ei)\u003c/b\u003e fail to reliably capture the disorder across scales,\u003csup\u003e30\u003c/sup\u003e \u003cb\u003eii)\u003c/b\u003e miss important symptoms,\u003csup\u003e31\u003c/sup\u003e \u003cb\u003eiii)\u003c/b\u003e differ substantially in their categorization of depressed patients into severity groups,\u003csup\u003e32\u003c/sup\u003e and \u003cb\u003eiv)\u003c/b\u003e are multidimensional, thus multiple constructs with inconsistent factor structures across scales are assessed.\u003csup\u003e\u003cspan additionalcitationids=\"CR34\" citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eAdditionally, the statistical method for combining information, which involves the use of algorithms or structured approaches, generally results in a 13% increase in accuracy over the clinical method, which relies on the clinician's judgment.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e,\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e Despite the strong evidence supporting the use of structured assessments and statistical methods, many clinicians continue to rely on less empirically supported practices. This disparity highlights a critical gap between research evidence and clinical practice.\u003c/p\u003e\n\u003ch3\u003eResearch on depression assessments in practice\u003c/h3\u003e\n\u003cp\u003eWith respect to clinical practice, research has examined how mental health professionals use diagnostic classification systems such as the \u003cem\u003eInternational Classification of Diseases\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/em\u003e\u003c/sup\u003e or the \u003cem\u003eDiagnostic and Statistical Manual of Mental Disorders (DSM)\u003c/em\u003e,\u003csup\u003e\u003cem\u003e38\u003c/em\u003e,39\u003c/sup\u003e what professionals\u0026rsquo; attitudes are toward standardized assessment,\u003csup\u003e40\u0026ndash;42\u003c/sup\u003e and how clinicians\u0026rsquo; perceptions of and approaches to the assessment process are related.\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e,\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\n\u003ch3\u003eStudy Objectives\u003c/h3\u003e\n\u003cp\u003eHow these methods are applied in the practices of clinicians is less known, and the available evidence suggests a discrepancy between guidelines on standards of practice and those of clinical practice.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e This study aims to gain insights into the extent to which empirically supported assessments are used by clinicians, with a particular focus on the assessment of depression, given that depression is highly prevalent and the leading cause of disability globally.\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e,\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e Specifically, we investigated (1) collection methods and instruments that clinicians use; (2) the approaches used to integrate and interpret these data (i.e., combination approaches), including clinical and statistical methods; (3) clinicians\u0026rsquo; attitudes toward standardized assessment tools and AI-based decision support; and (4) clinicians\u0026rsquo; beliefs about the accuracy of their assessments relative to other methods and other clinicians. These questions were explored across diverse clinical settings, professional roles, and countries, with the goal of understanding the extent to which empirically supported assessment practices are implemented in real-world clinical contexts and the factors influencing their adoption.\u003c/p\u003e \u003cp\u003eMethods\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eParticipants\u003c/h2\u003e \u003cp\u003eClinicians were recruited via convenience- and snowball sampling through several different online platforms such as LinkedIN, Facebook groups, emailing lists, workplace platforms (e.g., Slack channels for clinicians), and Prolific (an online platform where participants can be paid to take part). The inclusion criteria were as follows: clinicians who \u003cb\u003e1)\u003c/b\u003e are currently working or have past experience in clinical practice (physically or digitally) and \u003cb\u003e2)\u003c/b\u003e are assessing or treating depression in patients. Clinicians who had never worked with assessing and/or treating depression were excluded (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;98). The final sample consisted of 585 participating clinicians: 426 (73%) were female, 153 (26%) were male, and six (1%) indicated \u0026lsquo;other\u0026rsquo;. There were 162 (28%) clinicians from the United Kingdom, 148 (25%) from the United States, 127 (22%) from Sweden, 105 (18%) from the Netherlands, and 43 (7%) from other countries. Among the different occupations reported, psychologists (\u003cem\u003en\u0026thinsp;=\u003c/em\u003e\u0026thinsp;195; 33%) and nurses (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;172; 29%) made up the largest part of the sample. Other occupations were physicians (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;88; 15%), psychotherapists (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;86; 15%), psychiatrists (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;22; 4%), and 78 (13%) of the respondents selected \u0026lsquo;other' (see Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e in the Electronic Supplementary Material [ESM] which presents a cross-tabulation of clinician occupations by country). The clinicians reported a mean of 9.08 years (\u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;8.2) of experience assessing and/or treating depression. Almost half of the clinicians (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;275; 47%) reported assessing and/or treating depression \u003cem\u003edaily or several times a week\u003c/em\u003e. The other clinicians reported doing so \u003cem\u003eonce a week or less\u003c/em\u003e (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;118; 20%), \u003cem\u003eonce a month or less\u003c/em\u003e (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;95; 16%), or \u003cem\u003ehaving done this previously but not anymore\u003c/em\u003e (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;97; 17%).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eMaterial: Survey Content\u003c/h3\u003e\n\u003cp\u003eThe survey was designed to gather insights into clinicians' practices, attitudes, and experiences regarding depression assessment. It included questions on data collection and combination methods, attitudes toward standardized assessment and AI, self-perceived accuracy, clinical experience, and demographics. The six domains within the questionnaire were (1) \u003cem\u003eAssessment Methods\u003c/em\u003e, where clinicians reported on the data collection methods they use (e.g., unstructured, semi-structured, and structured interviews; rating scales), frequency of use, and specific instruments for both initial assessment and follow-up; (2) \u003cem\u003eData Combination Methods\u003c/em\u003e, where clinicians indicated whether they rely on clinical judgment, statistical models, or single data sources to integrate information; (3) \u003cem\u003eAssessment Attitudes\u003c/em\u003e, where clinicians were asked about their attitudes regarding the importance of standardized methods for diagnostic assessments, as well as their interest in AI decision support coupled with one of the randomly assigned descriptions (1\u0026thinsp;=\u0026thinsp;no information on the accuracy of AI, 2\u0026thinsp;=\u0026thinsp;assuming high accuracy of AI, 3\u0026thinsp;=\u0026thinsp;empirical data on accuracy of AI); (4) \u003cem\u003eSelf-Perceived Accuracy\u003c/em\u003e, where clinicians compared their accuracy against peers using similar or different methods, including AI and combinations of methods; (5) Clinical \u003cem\u003eExperience and Confidence\u003c/em\u003e, where clinicians reported their experience with depression assessments, their confidence levels, and time spent per assessment; and (6) \u003cem\u003eDemographics\u003c/em\u003e, including age, gender, country, profession, therapeutic approach, and employment sector. Detailed descriptions of the questions, response options, and randomized AI descriptions are provided on the ESM, along with the full questionnaire.\u003c/p\u003e\n\u003ch3\u003eProcedure\u003c/h3\u003e\n\u003cp\u003e \u003cstrong\u003eInformed consent\u003c/strong\u003e \u003cp\u003ewas obtained from all the clinicians at the start of the survey. They were informed about the nature of the study, that their participation was voluntary and anonymous, that they had the right to withdraw at any time without providing a reason, and that their anonymized answers would be shared according to open science practices. Clinicians completed the survey online, following the question sequence outlined in the Materials section. The clinicians were presented with one of the randomized conditions when asked about their attitudes toward AI decision support. However, owing to an issue with the survey tool, the assigned condition was only recorded for a subset of clinicians, specifically those who completed the survey through Prolific. After completing the survey, the clinicians were debriefed and given the opportunity to send feedback. The survey was available in English, Dutch, or Swedish and took, on average, 10.92 minutes (\u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;15.89; Med\u0026thinsp;=\u0026thinsp;7.82) to complete.\u003c/p\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analyses\u003c/h2\u003e \u003cp\u003eThe statistical analyses were performed in R (R Core Team, 2022; for details, see ESM). Descriptive statistics (frequencies [Freq], means [\u003cem\u003eM\u003c/em\u003e], standard deviations [\u003cem\u003eSD\u003c/em\u003e], medians [Med]) were calculated for each variable. Pearson correlations (\u003cem\u003er\u003c/em\u003e) were used to estimate the relationships between the variables, where the cut-off values for the different strengths were as follows: very weak or no correlation (\u0026plusmn;\u0026thinsp;.0-.2), weak (\u0026plusmn;\u0026thinsp;.3-.4), moderate (\u0026plusmn;\u0026thinsp;.5-.6), strong (\u0026plusmn;\u0026thinsp;.7-.8), and very strong (\u0026plusmn;\u0026thinsp;.9\u0026thinsp;\u0026minus;\u0026thinsp;1). Analysis of variance (ANOVA) was employed for the AI condition comparisons and the cross-cultural comparisons. All ANOVAs were calculated for unbalanced designs (Type III ANOVA). The assumptions of homogeneity of variances were checked via Levene\u0026rsquo;s test and were met. For significant main effects, post-hoc pairwise comparisons were conducted with Bonferroni correction to control for multiple testing. The \u003cem\u003ealpha\u003c/em\u003e-value was set to 0.05.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eDuration of evaluation\u003c/h2\u003e \u003cp\u003eThe self-reported time it took for clinicians to identify depression in a typical patient ranged from 3 minutes to 15 hours[1]\u003ca class=\"FNLink\" href=\"#Fn1\" id=\"#FNLinkFn1\"\u003e\u003c/a\u003e, with a median time of 60 minutes (\u003cem\u003eSD\u003c/em\u003e\u0026thinsp;=\u0026thinsp;108 minutes, Mean\u0026thinsp;=\u0026thinsp;87 minutes). Self-reported assessment time did not correlate significantly with the reported use of various data collection methods (see Supplementary Table S2 in the ESM for the correlations between all variables).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e\u003cem\u003eAim 1: Data collection methods\u003c/em\u003e\u003c/h2\u003e \u003cp\u003eThe clinicians reported which data collection method(s) they used (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). Most clinicians reported using unstructured interviews (83%) and/or rating scales (83%), followed by semistructured interviews (65%), structured interviews (40%), and/or \u0026lsquo;other\u0026rsquo; methods (25%).\u003c/p\u003e \u003cp\u003eThe clinicians also estimated the proportion of their patients for whom they used these different methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). On average, they reported using clinical interviews for 58% of their patients, followed by rating scales for 51% of their patients, and semistructured interviews for 34% of their patients.\u003c/p\u003e \u003cp\u003eMost clinicians (89%) reported using two or more methods across patients (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). Only 11% of the clinicians reported using a single method across all of their patients, of which the unstructured interviews were the most common method. The most commonly reported data collection method combinations were as follows: 1) Unstructured interviews and rating scales (18%); 2) Unstructured interviews, semistructured interviews, structured interviews, and rating scales (18%); and 3) Unstructured interviews, semistructured interviews, and rating scales (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;16%).\u003c/p\u003e \u003cp\u003eMost clinicians (84%) indicated performing follow-up assessments; similar patterns regarding the variety of data collection methods were observed among them. A more detailed description is available in the ESM.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;585\u003c/p\u003e \u003cp\u003e \u003cb\u003eSpecific data collection scales and (semi-)structured interviews.\u003c/b\u003e Clinicians could indicate which standardized tools and instruments they use. A list of the specific instruments and their corresponding figures is detailed in the ESM and the \u003cem\u003eMental Health Assessment Dashboard\u003c/em\u003e (see MHAD; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cwiebel.shinyapps.io/ClinicalPractices/\u003c/span\u003e\u003cspan address=\"https://cwiebel.shinyapps.io/ClinicalPractices/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), an interactive online tool that allows for further analysis and custom comparisons of the data.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eAim 2: Data combination methods\u003c/h2\u003e \u003cp\u003eThe vast majority (80%) of the clinicians reported combining the collected data with clinical experience (i.e., the clinical method; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Eleven percent reported using the statistical data combination method, and 6% reported not using any data combination method since they used only one source of information.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;585.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAim 3\u003c/b\u003e: \u003cb\u003eAttitudes towards assessment methods\u003c/b\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eAttitudes toward standardized assessment methods.\u003c/b\u003e More than half of the clinicians reported \u003cem\u003esomewhat\u003c/em\u003e to \u003cem\u003estrongly agree[ing]\u003c/em\u003e (76%) that it is important to use standardized assessment methods with high validity and reliability (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). A notable number of clinicians reported \u003cem\u003esomewhat-\u003c/em\u003e to \u003cem\u003estrongly disagree[ing]\u003c/em\u003e (17%) with this statement. Clinicians\u0026rsquo; ratings of the importance of using standardized assessments were correlated with their reported frequency of using rating scales (\u003cem\u003er\u003c/em\u003e = .20, \u003cem\u003ep\u003c/em\u003e \u0026lt;\u0026thinsp;.001) and structured interviews (\u003cem\u003er\u003c/em\u003e = .14, \u003cem\u003ep\u003c/em\u003e = .002) but negatively correlated with their reported frequency of using unstructured interviews (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;.13, \u003cem\u003ep\u003c/em\u003e = .002). Many clinicians have elaborated on their answers. A summary of these elaborations is presented in the SM.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;585.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAttitudes toward AI decision support.\u003c/b\u003e More clinicians reported to be \u003cem\u003eprobably\u003c/em\u003e or \u003cem\u003edefinitely\u003c/em\u003e interested in receiving AI decision support (43%) than \u003cem\u003eprobably not\u003c/em\u003e or \u003cem\u003edefinitely not\u003c/em\u003e (33%; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). A substantial number reported not knowing or having too little knowledge to take a stand (24%). Greater interest in using AI was correlated with positive agreement on the importance of using standardized assessments (\u003cem\u003er\u003c/em\u003e = .11, \u003cem\u003ep\u003c/em\u003e = .009) and with the reported frequency of using rating scales (\u003cem\u003er\u003c/em\u003e = .17, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001). The attitudes toward the use of AI decision support did not significantly differ between the AI information conditions (\u003cem\u003eF\u003c/em\u003e[2,324]\u0026thinsp;=\u0026thinsp;1.60, \u003cem\u003ep\u003c/em\u003e = .204, \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;359)[2]\u003ca class=\"FNLink\" href=\"#Fn2\" id=\"#FNLinkFn2\"\u003e\u003c/a\u003e. Many clinicians have also elaborated on their answers, which are presented on the ESM.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eAim 4: Beliefs in one\u0026rsquo;s confidence and accuracy\u003c/h2\u003e \u003cp\u003eA majority of the clinicians reported feeling \u003cem\u003econfident\u003c/em\u003e (35%), \u003cem\u003every confident\u003c/em\u003e (32%), or \u003cem\u003eextremely confident\u003c/em\u003e (10%) in identifying depression, whereas few reported feeling \u003cem\u003esomewhat confident\u003c/em\u003e (19%), \u003cem\u003ea little confident\u003c/em\u003e (4%), or \u003cem\u003enot at all confident\u003c/em\u003e (\u0026lt;\u0026thinsp;1%). Confidence in identifying depression was positively correlated with the reported assessment frequency of depression (\u003cem\u003er\u003c/em\u003e = .33, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), work experience with identifying depression (\u003cem\u003er\u003c/em\u003e = .21, \u003cem\u003ep\u003c/em\u003e \u0026lt; .001), and the reported frequency of using rating scales and semistructured interviews as data collection methods (\u003cem\u003er\u003c/em\u003e = .17, \u003cem\u003ep \u0026lt; .001\u003c/em\u003e and \u003cem\u003er = .11, p = .007\u003c/em\u003e). See the ESM for more details.\u003c/p\u003e \u003cp\u003eClinicians reported believing that their assessments (based on the method[s] they usually use for assessing depression) yield superior accuracy compared to those clinicians who use only one data collection method, but not with those clinicians who use a combination of data collection methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). The largest discrepancy between the clinicians\u0026rsquo; own estimated assessment accuracy compared to others was noted in comparison to those who relied only on rating scales as their single data collection source.\u003c/p\u003e \u003cp\u003eAmong the clinicians who reported using more than one method in their assessments (89%), most reported believing their accuracy to be similar to or slightly less than that of other clinicians using multiple methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003ec). However, 80% of clinicians rely on clinical experience to interpret results from multiple sources, and these clinicians reported believing their assessment to be more accurate than that of AI or statistical models (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e4\u003c/span\u003eb). In summary, while clinicians do not consider themselves more accurate than their peers do when a combination of methods is used, they do view their accuracy as superior to that of statistical models.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe error bars represent the standard error of the mean. \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;585.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eCross-cultural differences\u003c/h2\u003e \u003cp\u003eExploratory analyses comparing clinicians from the U.S., the Netherlands, Sweden, and the U.K. identified few small to moderate cross-cultural differences. Clinicians from the U.S. and the U.K. used unstructured interviews for the smallest percentage of patients, with averages of 51% and 46%, respectively. The finding that the vast majority of clinicians report combining collected data with clinical experience is consistent across countries. For a more exploratory cross-cultural analysis, see the ESM and the \u003cem\u003eMHAD\u003c/em\u003e. However, these findings should be interpreted with caution due to the heterogeneous nature of the samples, especially with respect to the occupations (e.g., psychologists, psychiatrists, general practitioners, nurses).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study examined the extent to which mental health professionals use empirically supported methods for assessing depression and their attitudes toward different methods. The results show that depression assessments are conducted by different professions with diverse educational backgrounds, using a wide range of methods and reporting large variations in the time it takes to assess, reflecting real-world practices and the diverse experiences patients encounter. The results show that (1) mental health professionals use a wide variety of data collection methods and that (2) they predominantly rely on the clinical method for data combination. (3) A total of 77% of the clinicians agreed on the importance of using standardized, valid, and reliable assessments, and 43% reported being interested in receiving AI decision support. (4) They are generally confident in their assessments but vary substantially in the reported time it takes to carry out a depression assessment. On average, clinicians report believing they are more accurate than AI or other clinicians are when a single data collection method is used. Interestingly, a wide variety of methods and practices were consistently found across all the examined countries, with only small significant cross-cultural differences. This highlights the paradox of consistent inconsistency in depression assessment practices on an international scale.\u003c/p\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eData collection methods\u003c/h2\u003e \u003cp\u003eClinicians report using different data collection methods with a wide range of combinations and instruments that vary greatly between individuals. Most clinicians (89%) reported using two or more methods, with unstructured interviews and rating scales being the most common. Although a significant portion of the clinicians (16%) did not conduct follow-up assessments, a similar pattern of diverse methods was observed among those who did. This aligns with previous studies highlighting the lack of use of outcome measures.\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e,\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e While unstructured interviews have shown low interrater reliability for diagnosing depression and are seldom used in research,\u003csup\u003e29,49\u003c/sup\u003e various rating scales for assessing depression differ significantly (i.e., they capture different symptoms of depression and only partly overlap\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e,\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e; and are often used interchangeably in research.\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e,\u003cspan additionalcitationids=\"CR50\" citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e This variation in data collection methods may be problematic for providing equal, standardized, high-quality care, especially since research shows that different methods and instruments yield varying results. The absence of a single industry standard for diagnostic manuals further contributes to inconsistencies in assessments and diagnoses.\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eData combination methods\u003c/h2\u003e \u003cp\u003eWhile approximately one in ten clinicians reported using the statistical method to form their final assessment, 80% reported interpreting and weighing the collected data on the basis of their clinical experience and judgment. Despite extensive research showing that the statistical method for combining information is as good as or superior to the clinical judgment of an expert clinician,\u003csup\u003e6,13,24,36\u003c/sup\u003e the majority of final assessments still rely on the clinical method. As early as 35 years ago, Dawes et al.\u003csup\u003e14\u003c/sup\u003e argued for the routine use of statistical methods in clinical assessment and decision-making. Grove and Meehl\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e addressed this issue, concluding that using the less accurate of two assessment methods \u0026ldquo;is not only unscientific and irrational, it is unethical\u0026rdquo; (\u003cem\u003ep\u003c/em\u003e. 323). However, possible explanations for the continued reliance on clinical data combination might be the lack of easily implemented, user-friendly, and validated statistical models that comply with the required regulations (e.g., CE-marking and FDA approval). Clinical validation and readiness for implementation should be weighed heavily before employing statistical models since, for example, flawed algorithms trained on small or biased training data can have adverse effects.\u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e This concern is further amplified by widespread apprehensions about the safety and validity of AI models, as highlighted by clinicians in this survey.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eAttitudes toward standardized assessment\u003c/h2\u003e \u003cp\u003eA sizable group of clinicians, approximately one-fifth, report that they \u003cem\u003esomewhat\u003c/em\u003e- to \u003cem\u003estrongly disagree\u003c/em\u003e with the importance of using standardized assessment methods that have been demonstrated in research to have high validity and reliability \u0026mdash; which we expected to find a unanimous agreement on. Even though 76% of the clinicians reported that they \u003cem\u003esomewhat-\u003c/em\u003e to \u003cem\u003estrongly agree\u003c/em\u003e with the statement, the importance they assigned to standardized assessment methods did not affect the final decision for most clinicians, who still use a combination of clinical data methods. This corresponds to other research examining clinicians\u0026rsquo; attitudes, which have shown a mostly, but not completely, positive attitude towards standardized assessment methods: Clinicians seem to be positive about the quality of standardized assessment methods but less positive about their usefulness in practice, and the advantages of standardized assessment over clinical judgment.\u003csup\u003e\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e To enable better implementation, the need to overcome specific barriers and issues, such as adequate training and practicality, is highlighted.\u003csup\u003e\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e,\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eAttitudes toward AI\u003c/h2\u003e \u003cp\u003eClinicians\u0026rsquo; interest in AI decision support for depression assessment is mixed: 43% express interest, 33% show disinterest, and 23% remain undecided due to insufficient information. This highlights the need to update clinicians on AI advancements and benefits and involve them in algorithm development and validation. Key concerns, including accountability, transparency, and privacy, must be addressed.\u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e Interest in AI support correlates with the use of rating scales and positive attitudes toward standardized methods, suggesting that clinicians familiar with these practices could facilitate AI integration in clinical settings. However, enhanced education on AI and statistical models is essential. For example, a study on psychiatrists\u0026rsquo; views of generative AI found widespread use for administrative tasks but skepticism toward its active role in patient care.\u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e The lack of differences in attitudes across AI descriptions and diverse responses in open comments further emphasize the need for comprehensive educational initiatives to foster an understanding of AI\u0026rsquo;s capabilities, limitations, and ethical integration into mental health practices.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eClinicians Assessment Accuracy\u003c/h2\u003e \u003cp\u003eStrikingly, a majority of clinicians who rely on a combination of clinical data believe that their assessments are more accurate than those derived from a statistical method or AI, which contradicts findings from previous research.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e,\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e The clinicians ranked the accuracy of their assessments as superior to those conducted by peers who use only a single method (particularly rating scales alone) but not compared with those using a combination of methods. These findings suggest a general belief among clinicians in the superiority of employing multiple methods for data collection and combining data via a clinical method. The belief in the clinical method could be explained by a more general tendency of individuals to overestimate their cognitive abilities.\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e However, clinicians\u0026rsquo; confidence and accuracy beliefs have been shown to be poor indicators of their assessment accuracy,\u003csup\u003e20,23,56\u003c/sup\u003e which is troubling given the impact it may have on patients and their course of treatment.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eThe results are constrained by certain limitations, yet they indicate several directions for future research to further our understanding of depression assessment practices and facilitate the adoption of empirically supported assessment methods. Given the study\u0026rsquo;s broad focus on depression assessments in real-world clinical settings, the sample comprises a wide range of occupations, including psychologists, psychotherapists, psychiatrists, and general practitioners. However, this makes it challenging to understand the experiences of specific occupations; therefore, future research could focus on single professional groups to gain deeper insights into their unique practices and challenges. Similarly, the goals and methods of clinicians\u0026rsquo; assessments can differ significantly both between and within countries, occupations, and clinics, all of which warrant further investigation.\u003c/p\u003e \u003cp\u003e This study included only Western countries, each with distinct guidelines, rules, and regulations governing mental health practices. The distribution of occupations also varies across countries, which necessitates caution when interpreting cross-country differences. Future research should expand to include non-Western countries to gain insights into how mental health practices are carried out in diverse cultural and regulatory contexts.\u003c/p\u003e \u003cp\u003eAdditionally, this study specifically examined the assessment of depression, so the findings may not be generalizable to other mental health conditions. The use of a self-report survey format introduces the potential for recall bias and may reflect clinicians\u0026rsquo; perceptions of best practices rather than their actual behaviors. As a result, the data could overestimate the extent to which empirically supported assessment methods are employed in practice. Finally, online recruitment and sampling may have restricted participation to specific subgroups. The sampling strategy is also likely to favor digitally active clinicians, which should be considered when interpreting the findings. In addition, AI-related randomization data were missing for some participants, which reduced the statistical power. Future research could benefit from random sampling directly from clinical settings or from observational studies of audits of actual practices, etc.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003ePotential implications and future research\u003c/h2\u003e \u003cp\u003eA better understanding of current depression assessment practices can provide clinicians, policymakers, and researchers with a foundation for improving healthcare quality. Understanding current practices can support clinicians, regulatory bodies, and researchers in joining forces to lead changes and update guidelines and policies.\u003c/p\u003e \u003cp\u003eThe Mental Health Assessment Dashboard can inform discussions on the current (\u003cem\u003elimited\u003c/em\u003e) use of empirically supported assessments and any future integration of AI in clinical settings. Furthermore, there is a need for more research regarding the potential risks of current depression assessment practices and their implications for patients and the challenges for clinicians in embracing newer empirically supported methods.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e There is a lack of research on how clinicians combine assessment methods and the potential shortcomings associated with the clinical combination method.\u003c/p\u003e \u003cp\u003e The variability in assessors and assessment methods reflects real-world practice and highlights the need for clearer guidelines and training to ensure equitable and consistent standardization. In other words, the outcome of a depression assessment for depression should not depend on where, when, or by whom the patient is assessed.\u003c/p\u003e \u003cp\u003eAdditionally, many clinicians\u0026rsquo; concerns about AI should be addressed, including biases and ethical issues, as well as how they can work with AI. Transparent communication and education about the capabilities and limitations of AI, along with evidence from robust studies, can potentially build trust and acceptance among clinicians.\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003ePrevious research\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e has typically studied depression assessment methods in research settings or focused on attitudes in clinical settings\u003csup\u003e\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e rather than actual assessment practices in different clinical practices. Our results highlight heterogeneity in depression assessments: Many different professionals conduct depression assessments in clinical practice (e.g., psychiatrists, psychologists, general practitioners, nurses); employ a wide variety of methods (ranging from unstructured to structured approaches using numerous tools) in different combinations, and report a great difference in the time it takes to carry out the assessment.\u003c/p\u003e \u003cp\u003eThe results reveal a pattern of skepticism toward AI, underestimation of statistical models, and overestimation of one\u0026rsquo;s own accuracy, which aligns with findings from previous studies. While clinicians express confidence in their ability to assess depression, the widespread use of unstructured interviews alongside standardized tools, coupled with a predominant reliance on clinical judgment over statistical models, suggests a disconnect between best practices supported by research and everyday clinical practice. Adopting empirically supported practices, including standardized, valid, and reliable data collection and combination assessment methods, is essential to synthesize information systematically and improve the quality of the data used in clinical judgment. This alignment between clinical practice and research evidence would enhance patient care while allowing clinicians to consider individual factors. However, achieving this requires harmonizing clinical judgment with data-driven approaches, as well as aligning manuals and guidelines with empirical research. Future initiatives should focus on developing accessible, practical, validated tools and training programs that bridge the gap between research and practice, fostering integration and consistency in depression assessment.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDSM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDiagnostic and Statistical Manual of Mental Disorders\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eESM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eElectronic Supplementary Material\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInternational Classification of Diseases\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLLM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLarge Language Model \u003cem\u003e(only include if mentioned earlier; I didn\u0026rsquo;t sit explicitly)\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMHAD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMental Health Assessment Dashboard\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eU.K.\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUnited Kingdom\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eU.S.\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eWHO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eWorld Health Organization\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eThe study recieved ethical approval from Leiden University (2022-08-10-M.L. Molendijk-V5-4037) and was exempt from ethical approval at the University of Pennsylvania (United States). The study is deemed exempt from requiring ethical approval according to Swedish Law (s§3-4 of the Act [2003:460] on the ethical review of research involving humans in Sweden). \u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003eAvailability of data and materials\u003c/h2\u003e\n\u003cp\u003eThe datasets generated and/or analysed during the current study are available at https://osf.io/pw8mk/. The code for the \u003cem\u003eMental Health Assessment Dashboard\u003c/em\u003e is also shared here in case the app no longer can be hosted (https://cwiebel.shinyapps.io/ClinicalPractices/).\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eO.N.E. Kjell and K. Kjell co-founded a start-up that uses computational language assessments to assess mental health problems. E.C. Stade has received fees for advising a start-up that uses computational language assessment to measure psychological constructs. The other authors report having no conflicts of interest with respect to the contents, authorship, or publication of this article.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eO.N.E. Kjell, K. Kjell and V.C. Eijsbroek were funded by FORTE (2022-01022).\u003c/p\u003e\n\u003ch2\u003eAuthors' contributions\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eConceptualization:\u003c/strong\u003e Ideas; formulation or evolution of overarching research goals and aims.\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData curation: \u003c/strong\u003eManagement activities to annotate (produce metadata), scrub data and maintain research data (including software code, where it is necessary for interpreting the data itself) for initial use and later re-use.\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFormal analysis:\u003c/strong\u003e Application of statistical, mathematical, computational, or other formal techniques to analyze or synthesize study data.\u003c/p\u003e\n\u003cp\u003eClara Wiebel, Veerle C. Eijsbroek, Damon Navandi, and Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding acquisition: \u003c/strong\u003eAcquisition of the financial support for the project leading to this publication.\u003c/p\u003e\n\u003cp\u003eKatarina Kjell, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInvestigation: \u003c/strong\u003eConducting a research and investigation process, specifically performing the experiments, or data/evidence collection.\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethodology: \u003c/strong\u003eDevelopment or design of methodology; creation of models.\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProject administration: \u003c/strong\u003eManagement and coordination responsibility for the research activity planning and execution.\u003c/p\u003e\n\u003cp\u003eKatarina Kjell, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResources: \u003c/strong\u003eProvision of study materials, reagents, materials, patients, laboratory samples, animals, instrumentation, computing resources, or other analysis tools.\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Katarina Kjell, Marc L. Molendijk, Elizabeth C. Stade, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSoftware: \u003c/strong\u003eProgramming, software development; designing computer programs; implementation of the computer code and supporting algorithms; testing of existing code components.\u003c/p\u003e\n\u003cp\u003eClara Wiebel\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupervision:\u003c/strong\u003e Oversight and leadership responsibility for the research activity planning and execution, including mentorship external to the core team.\u003c/p\u003e\n\u003cp\u003eOscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eValidation: \u003c/strong\u003eVerification, whether as a part of the activity or separate, of the overall replication/reproducibility of results/experiments and other research outputs.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eVisualization:\u003c/strong\u003e Preparation, creation and/or presentation of the published work, specifically visualization/data presentation.\u003c/p\u003e\n\u003cp\u003eClara Wiebel, Veerle C. Eijsbroek, Damon Navandi, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWriting—original draft: \u003c/strong\u003ePreparation, creation and/or presentation of the published work, specifically writing the initial draft (including substantive translation).\u003c/p\u003e\n\u003cp\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel, Oscar N. E. Kjell\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWriting—review and editing: \u003c/strong\u003ePreparation, creation and/or presentation of the published work by those from the original research group, specifically critical review, commentary or revision: including pre- or post-publication stages.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDamon Navandi, Veerle C. Eijsbroek, Clara Wiebel\u003c/strong\u003e\u003c/p\u003e\n\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eShen H, Zhang L, Xu C, Zhu J, Chen M, Fang Y. Analysis of Misdiagnosis of Bipolar Disorder in An Outpatient Setting. Shanghai Archives of Psychiatry [Internet]. 2018;30(2):93\u0026ndash;101. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC5936046/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5936046/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStahnke B. A systematic review of misdiagnosis in those with obsessive-compulsive disorder. Journal of Affective Disorders Reports [Internet]. 2021;6(1):100231. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.sciencedirect.com/science/article/pii/S2666915321001578\u003c/span\u003e\u003cspan address=\"https://www.sciencedirect.com/science/article/pii/S2666915321001578\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFried EI, Proppert RKK, Rieble CL. Building an Early Warning System for Depression: Rationale, Objectives, and Methods of the WARN-D Study. Clinical Psychology in Europe [Internet]. 2023;5(3):1\u0026ndash;25. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cpe.psychopen.eu/index.php/cpe/article/view/10075\u003c/span\u003e\u003cspan address=\"https://cpe.psychopen.eu/index.php/cpe/article/view/10075\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEijsbroek VC, Katarina Kjell, Schwartz A, Boehnke JR, Fried EI, Klein DN et al. The LEADING Guideline. Reporting Standards for Expert Panel, Best-Estimate Diagnosis, and Longitudinal Expert All Data (LEAD) Studies. medRxiv (Cold Spring Harbor Laboratory). 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHunsley J, Mash EJ. Evidence-based assessment. Annual review of clinical psychology [Internet]. 2007;3:29\u0026ndash;51. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/pubmed/17716047\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/pubmed/17716047\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeehl PE. Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. Minneapolis: University of Minnesota Press; 1954.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePettersson A, Bostr\u0026ouml;m KB, Gustavsson P, Ekselius L. Which instruments to support diagnosis of depression have sufficient accuracy? A systematic review. Nord J Psychiatry. 2015;69(7):497\u0026ndash;508.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCook JR, Hausman EM, Jensen-Doss A, Hawley KM. Assessment Practices of Child Clinicians. Assessment. 2016;24(2):210\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGraham S, Depp C, Lee EE, Nebeker C, Tu X, Kim HC, et al. Artificial Intelligence for Mental Health and Mental Illnesses: an Overview. Curr Psychiatry Rep. 2019;21(11):116.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOscar NE, Kjell K, Kjell H. Andrew Schwartz. Beyond Rating Scales: With Targeted Evaluation, Language Models are Poised for Psychological Assessment. Psychiatry Res. 2023;333:115667\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee EE, Torous J, De Choudhury M, Depp CA, Graham SA, Kim HC, et al. Artificial Intelligence for Mental Healthcare: Clinical Applications, Barriers, Facilitators, and Artificial Wisdom. Biol Psychiatry: Cogn Neurosci Neuroimaging. 2021;6(9):856\u0026ndash;64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMueller AE, Segal DL. Structured versus Semistructured versus Unstructured Interviews. Encyclopedia Clin Psychol. 2015;1(7):1\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\u0026AElig;gisd\u0026oacute;ttir S, White MJ, Spengler PM, Maugherman AS, Anderson LA, Cook RS, et al. The Meta-Analysis of Clinical Judgment Project: Fifty-Six Years of Accumulated Research on Clinical Versus Statistical Prediction. Couns Psychol. 2006;34(3):341\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDawes R, Faust D, Meehl P. Clinical versus actuarial judgment. Science [Internet]. 1989;243(4899):1668\u0026ndash;74. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://meehl.umn.edu/sites/meehl.dl.umn.edu/files/138cstixdawesfaustmeehl.pdf\u003c/span\u003e\u003cspan address=\"http://meehl.umn.edu/sites/meehl.dl.umn.edu/files/138cstixdawesfaustmeehl.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeRubeis RJ. The history, current status, and possible future of precision mental health. Behav Res Ther. 2019;123:103506.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShatte ABR, Hutchinson DM, Teague SJ. Machine learning in mental health: a scoping review of methods and applications. Psychological Medicine [Internet]. 2019;49(09):1426\u0026ndash;48. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cambridge.org/core/journals/psychological-medicine/article/abs/machine-learning-in-mental-health-a-scoping-review-of-methods-and-applications/0B70B1C827B3A4604C1C01026049F7D9\u003c/span\u003e\u003cspan address=\"https://www.cambridge.org/core/journals/psychological-medicine/article/abs/machine-learning-in-mental-health-a-scoping-review-of-methods-and-applications/0B70B1C827B3A4604C1C01026049F7D9\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBennett CC, Hauser K. Artificial intelligence framework for simulating clinical decision-making: A Markov decision process approach. Artif Intell Med. 2013;57(1):9\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEgon Brunswik, Cooksey RW. The Conceptual Framework of Psychology [1952]. 2001;225\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHammond KR, Hursch CJ, Todd FJ. Analyzing the components of clinical inference. Psychol Rev. 1964;71(6):438\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarelaia N, Hogarth RM. Determinants of linear judgment: A meta-analysis of lens model studies. Psychol Bull. 2008;134(3):404\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrove WM, Meehl PE. Comparative efficiency of informal (subjective, impressionistic) and formal (mechanical, algorithmic) prediction procedures: The clinical-statistical controversy. Psychol Public Policy Law. 1996;2(2):293\u0026ndash;323.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBowes SM, Ammirati RJ, Costello TH, Basterfield C, Lilienfeld SO. Cognitive biases, heuristics, and logical fallacies in clinical practice: A brief field guide for practicing clinicians and supervisors. Prof Psychology: Res Pract. 2020;51(5):435\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaposnik G, Redelmeier D, Ruff CC, Tobler PN. Cognitive biases associated with medical decisions: a systematic review. BMC Med Inf Decis Mak. 2016;16(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarb HN, Wood JM. Methodological advances in statistical prediction. Psychol Assess. 2019;31(12):1456\u0026ndash;66.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollins FS, Varmus H. A New Initiative on Precision Medicine. N Engl J Med. 2015;372(9):793\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDavenport T, Kalakota R. The Potential for Artificial Intelligence in Healthcare. Future Healthcare Journal [Internet]. 2019;6(2):94\u0026ndash;8. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pmc.ncbi.nlm.nih.gov/articles/PMC6616181/\u003c/span\u003e\u003cspan address=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC6616181/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAli GC, Ryan G, De Silva MJ. Validated Screening Tools for Common Mental Disorders in Low and Middle Income Countries: A Systematic Review. Burns JK, editor. PLOS ONE. 2016;11(6):e0156939.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreedvelt JJF, Zamperoni V, South E, Uphoff EP, Gilbody S, Bockting CLH, et al. A systematic review of mental health measurement scales for evaluating the effects of mental health prevention interventions. Eur J Pub Health. 2020;30(3):510\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiller PR, Dasher R, Collins R, Griffiths P, Brown F. Inpatient diagnostic assessments: 1. Accuracy of structured vs. unstructured interviews. Psychiatry Res. 2001;105(3):255\u0026ndash;64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFried EI. The 52 symptoms of major depression: Lack of content overlap among seven common depression scales. J Affect Disord. 2017;208(208):191\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChevance A, Ravaud P, Tomlinson A, Le Berre C, Teufer B, Touboul S, et al. Identifying outcomes for depression that matter to patients, informal caregivers, and health-care professionals: qualitative content analysis of a large international online survey. Lancet Psychiatry. 2020;7(8):692\u0026ndash;702.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZimmerman M, Martinez JH, Friedman M, Boerescu DA, Attiullah N, Toba C. How Can We Use Depression Severity to Guide Treatment Selection When Measures of Depression Categorize Patients Differently? J Clin Psychiatry. 2012;73(10):1287\u0026ndash;91.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShafer AB. Meta-analysis of the factor structures of four depression questionnaires: Beck, CES-D, Hamilton, and Zung. J Clin Psychol. 2005;62(1):123\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Loo HM, de Jonge P, Romeijn JW, Kessler RC, Schoevers RA. Data-driven subtypes of major depressive disorder: a systematic review. BMC Med. 2012;10(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFried EI, van Borkulo CD, Epskamp S, Schoevers RA, Tuerlinckx F, Borsboom D. Measuring depression over time.. Or not? Lack of unidimensionality and longitudinal measurement invariance in four common rating scales of depression. Psychol Assess. 2016;28(11):1354\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrove WM, Zald DH, Lebow BS, Snitz BE, Nelson C. Clinical versus mechanical prediction: A meta-analysis. Psychol Assess. 2000;12(1):19\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eICD-11 [Internet]. icd.who.int. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://icd.who.int\u003c/span\u003e\u003cspan address=\"https://icd.who.int\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmerican Psychiatric Association. Diagnostic and statistical manual of mental disorders. Diagnostic and Statistical Manual of Mental Disorders. 5th ed. 2013;5(5).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFirst MB, Rebello TJ, Keeley JW, Bhargava R, Dai Y, Kulygina M, et al. Do mental health professionals use diagnostic classifications the way we think they do? A global survey. World Psychiatry. 2018;17(2):187\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDanielson M, M\u0026aring;nsdotter A, Fransson E, Dalsgaard S, Larsson J-O. Clinicians\u0026rsquo; attitudes toward standardized assessment and diagnosis within child and adolescent psychiatry. Child Adolesc Psychiatry Mental Health. 2019;13(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJensen-Doss A, Hawley KM. Understanding Barriers to Evidence-Based Assessment: Clinician Attitudes Toward Standardized Assessment Tools. Journal of Clinical Child \u0026amp; Adolescent Psychology [Internet]. 2010;39(6):885\u0026ndash;96. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC3058768/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3058768/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJensen-Doss A, Hawley KM. Understanding Clinicians\u0026rsquo; Diagnostic Practices: Attitudes Toward the Utility of Diagnosis and Standardized Diagnostic Tools. Administration and Policy in Mental Health and Mental Health Services Research [Internet]. 2011;38(6):476\u0026ndash;85. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC6114089/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6114089/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBohman B. Clinicians\u0026rsquo; perceptions and practices of diagnostic assessment in psychiatric services. BMC Psychiatry. 2023;23(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho E, Tugendrajch SK, Marriott BR, Hawley KM. Evidence-Based Assessment in Routine Mental Health Services for Youths. Psychiatric Serv. 2020;72(3):\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003eappi.ps.2019005\u003c/span\u003e\u003cspan address=\"http://appi.ps.2019005\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMathers CD, Loncar D. Projections of Global Mortality and Burden of Disease from 2002 to 2030. Samet J, editor. PLoS Medicine [Internet]. 2006;3(11):e442. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/17132052/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/17132052/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrevalence. and variability of current depressive disorder in 27 European countries: a population-based study. The Lancet Public Health [Internet]. 2021; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.sciencedirect.com/science/article/pii/S2468266721000475#bib9\u003c/span\u003e\u003cspan address=\"https://www.sciencedirect.com/science/article/pii/S2468266721000475#bib9\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoswell JF, Kraus DR, Miller SD, Lambert MJ. Implementing routine outcome monitoring in clinical practice: Benefits, challenges, and solutions. Psychother Res. 2013;25(1):6\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGilbody SM, House AO, Sheldon TA. Psychiatrists in the UK do not use outcomes measures. Br J Psychiatry. 2002;180(2):101\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRegier DA, Narrow WE, Clarke DE, Kraemer HC, Kuramoto SJ, Kuhl EA, et al. DSM-5 Field Trials in the United States and Canada, Part II: Test-Retest Reliability of Selected Categorical Diagnoses. Am J Psychiatry. 2013;170(1):59\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNewson JJ, Hunter D, Thiagarajan TC. The heterogeneity of mental health assessment. Frontiers in Psychiatry [Internet]. 2020;11(76):1\u0026ndash;24. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.frontiersin.org/articles/\u003c/span\u003e\u003cspan address=\"https://www.frontiersin.org/articles/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fpsyt.2020.00076/full\u003c/span\u003e\u003cspan address=\"10.3389/fpsyt.2020.00076/full\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSantor DA, Gregus M, Welch A. FOCUS ARTICLE: Eight Decades of Measurement in Depression. Measurement: Interdisciplinary Research \u0026amp; Perspective. 2006;4(3):135\u0026ndash;55.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTopol EJ. High-performance medicine: the Convergence of Human and Artificial Intelligence. Nature Medicine [Internet]. 2019;25(1):44\u0026ndash;56. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.nature.com/articles/s41591-018-0300-7\u003c/span\u003e\u003cspan address=\"https://www.nature.com/articles/s41591-018-0300-7\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeslie D, systems in the public sector Dr David Leslie Public Policy Programme. Understanding artificial intelligence ethics and safety A guide for the responsible design and implementation of AI. Understanding artificial intelligence ethics and safety [Internet]. 2019; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.turing.ac.uk/sites/default/files/2019-06/understanding_artificial_intelligence_ethics_and_safety.pdf\u003c/span\u003e\u003cspan address=\"https://www.turing.ac.uk/sites/default/files/2019-06/understanding_artificial_intelligence_ethics_and_safety.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlease C, Worthen A, Torous J. Psychiatrists\u0026rsquo; Experiences and Opinions of Generative Artificial Intelligence in Mental Healthcare: An Online Mixed Methods Survey. Psychiatry Res. 2024;333:115724\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeck PR, Simons DJ, Chabris CF. 65% of Americans believe they are above average in intelligence: Results of two nationally representative surveys. van Amelsvoort T, editor. PLOS ONE. 2018;13(7):e0200103.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiller DJ, Spengler ES, Spengler PM. A meta-analysis of confidence and judgment accuracy in clinical decision making. J Couns Psychol. 2015;62(4):553\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBetancourt TS, Chambers DA. Optimizing an Era of Global Mental Health Implementation Science. JAMA Psychiatry. 2016;73(2):99.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMisra R, Keane PA, Hogg HDJ. How should we train clinicians for artificial intelligence in healthcare? Future Healthc J. 2024;11(3):100162.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Footnotes","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eNine cases with a reported time greater than the mean plus 3 standard deviations were seen as outliers and thus removed from the time analysis. Their reported assessment times ranged from 20 to 60 hours, respectively. Additionally, seven cases reported 0 minutes; these were removed because it is possible that they do not assess depression at all.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eThe lower number of clinicians in this analysis is due to the randomisation condition not being saved for all clinicians.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-psychology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"psyo","sideBox":"Learn more about [BMC Psychology](http://bmcpsychology.biomedcentral.com/)","snPcode":"","submissionUrl":"","title":"BMC Psychology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"clinical assessment, depression, structured methods, statistical models, artificial intelligence","lastPublishedDoi":"10.21203/rs.3.rs-8951240/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8951240/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurate psychological evaluation is crucial for early detection, intervention and evaluation of mental health disorders. The validity and reliability of psychological evaluation methods have been systematically researched for many decades. However, assessment methods used by mental health professionals in clinical settings have rarely been empirically documented. This study examines the extent to which empirically supported assessments are used in clinical settings and clinicians\u0026rsquo; attitudes toward these practices. Clinicians (\u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;585) from the U.S. (25%), the U.K. (28%), the Netherlands (18%), Sweden (22%), and other countries (7%) completed an online survey reporting their depression assessment methods. Clinicians reported using many different \u003cem\u003edata collection methods\u003c/em\u003e (e.g., unstructured clinical interviews [83%], rating scales [83%]) in various combinations. A majority (80%) primarily employed clinical methods, as opposed to statistical \u003cem\u003edata combination methods\u003c/em\u003e (11%). Although they were confident in their assessments, the time to carry out assessments varied widely (\u003cem\u003eM\u003c/em\u003e\u0026thinsp;=\u0026thinsp;87 [\u003cem\u003eSD\u0026thinsp;=\u003c/em\u003e\u0026thinsp;108] minutes). Seventy-seven percent agreed on the importance of standardized, valid, and reliable assessments, and 43% reported interest in receiving AI decision-support. We did not find many significant cross-cultural differences. A majority of clinicians have not fully adopted standardized empirically supported assessment practices. Predominant reliance on clinical judgment suggests a disconnect between best practices supported by research and everyday clinical practice.\u003c/p\u003e","manuscriptTitle":"Clinicians’ Practices and Attitudes Toward Depression Assessment: A Cross-Sectional Survey Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-04 09:00:08","doi":"10.21203/rs.3.rs-8951240/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-23T08:32:39+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-19T17:44:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"112134730657901396563315757813958980793","date":"2026-03-19T13:53:21+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-09T15:08:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"31855867118683275702147632381034439701","date":"2026-03-02T00:10:22+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"201168911739626933382733310364127064371","date":"2026-02-27T17:30:55+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-27T15:19:12+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-26T15:00:59+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-26T01:44:45+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-26T01:43:23+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Psychology","date":"2026-02-23T23:56:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-psychology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"psyo","sideBox":"Learn more about [BMC Psychology](http://bmcpsychology.biomedcentral.com/)","snPcode":"","submissionUrl":"","title":"BMC Psychology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8dd10795-ec9e-405f-bdf3-612bafe1b593","owner":[],"postedDate":"March 4th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-27T20:54:20+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-04 09:00:08","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8951240","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8951240","identity":"rs-8951240","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.