Generative artificial intelligence in mental health: A preliminary study on automating materials development for cognitive bias modification.

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract People with depression tend to interpret ambiguous events in a negative biased direction, which contributes to symptomatology. Cognitive bias modification-interpretation (CBM-I) is a digital therapeutic that targets negative interpretation bias using text-based scenarios. CBM-I offers greater flexibility in combating depressive-related issues, but the development of its training materials can be costly. In the present study, we used Generative AI to produce CBM-I training materials and compared them with human-generated materials. The aim is to examine whether AI-generated materials are equivalent to human-generated materials, as per a set of pre-defined criteria designed to capture common experiences of individuals with depression. We followed the typical CBM-I materials development procedure, first creating raw items and then adapting them into standard CBM-I format. We compared participants’ ratings of 100 raw scenarios and 100 CBM-I scenarios, half of which were created by Copilot and half created by people with depression. Living/lived experts of depression (N = 30) rated raw items, and another 30 experts rated CBM-I items on readability, relevance, and severity of scenarios as related to depression. With the exception of severity ratings, results revealed that ratings of human-generated and AI-generated scenarios were statistically non-equivalent. The differences in the overall actual mean ratings, however, were small (range 0.03–0.41); the overall direction of ratings between AI-and human-generated scenarios were the same and consistent with the scenarios’ emotional content. Interpretation of the data and future implications of outsourcing AI in CBM-I materials production are discussed.
Full text 141,475 characters · extracted from preprint-html · click to expand
Generative artificial intelligence in mental health: A preliminary study on automating materials development for cognitive bias modification. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Generative artificial intelligence in mental health: A preliminary study on automating materials development for cognitive bias modification. Che-Wei Hsu, Alex Robbins, Tiana Cartwright This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7530420/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract People with depression tend to interpret ambiguous events in a negative biased direction, which contributes to symptomatology. Cognitive bias modification-interpretation (CBM-I) is a digital therapeutic that targets negative interpretation bias using text-based scenarios. CBM-I offers greater flexibility in combating depressive-related issues, but the development of its training materials can be costly. In the present study, we used Generative AI to produce CBM-I training materials and compared them with human-generated materials. The aim is to examine whether AI-generated materials are equivalent to human-generated materials, as per a set of pre-defined criteria designed to capture common experiences of individuals with depression. We followed the typical CBM-I materials development procedure, first creating raw items and then adapting them into standard CBM-I format. We compared participants’ ratings of 100 raw scenarios and 100 CBM-I scenarios, half of which were created by Copilot and half created by people with depression. Living/lived experts of depression (N = 30) rated raw items, and another 30 experts rated CBM-I items on readability, relevance, and severity of scenarios as related to depression. With the exception of severity ratings, results revealed that ratings of human-generated and AI-generated scenarios were statistically non-equivalent. The differences in the overall actual mean ratings, however, were small (range 0.03–0.41); the overall direction of ratings between AI-and human-generated scenarios were the same and consistent with the scenarios’ emotional content. Interpretation of the data and future implications of outsourcing AI in CBM-I materials production are discussed. Psychology cognitive bias modification interpretation bias depression ChatGPT generative AI chatbot Copilot digital mental health Figures Figure 1 Figure 2 Figure 3 Introduction According to cognitive models of psychological disorders, interpretation bias is thought to contribute to the formation and maintenance of various mental health concerns (Beck, 1976; Garety et al., 2001). Interpretation bias, in this context can be defined as the tendency to interpret ambiguous situations with a negative or threatening connotation. Across many studies, researchers have found compelling evidence of interpretation bias among people who experience various mental health issues such as depression (Everaert et al., 2017), anxiety (Chen et al., 2020), and paranoia (Savulich et al., 2015). More importantly from an intervention standpoint, it has been well-documented that interpretation bias can be modified to reduce mental health symptoms (Li et al., 2022; Yiend et al., 2013). Cognitive bias modification-Interpretation (CBM-I) is a class of evidence-based cognitive therapeutic designed to directly target damaging interpretation bias to improve symptomology. There are several features of CBM-I that may make it an appealing alternative to traditional psychological interventions such as cognitive behavioural therapy (CBT). First, CBM-I is self-administered via a mobile phone app or a web platform, which means the treatment is highly accessible, scalable, and discrete, making it easier to disseminate to a wider population. It is also designed to be a low-cost treatment (Ha & Kim, 2020), with cost referring to time, personnel, and financial expenses. Additionally, CBM-I can be simpler to implement than CBT and offers a more targeted approach to treatment, which may enhance its clinical application. Despite these advantages over other therapeutics, developing training materials for CBM-I can be challenging and costly, thereby limiting its widespread utility and availability in clinical settings. The overarching goal of the present study is to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials. In this context, we define equivalence in a narrow sense, basing it on a set of pre-defined criteria designed to capture common experiences of individuals with depression. CBM-I adopts findings from basic research of interpretation bias and its association with depression. More specifically, people with depression tend to interpret ambiguous everyday events in a more negative direction (Rude et al., 2003); using text-based excerpts of such events could successfully evoke and modify people’s negative interpretation bias to reduce depression (Li et al., 2022; Mathews & Mackintosh, 2000; Yiend et al., 2013). CBM-I employs a set of ambiguous scenarios designed to elicit multiple pathways of interpretation. By using a word task, CBM-I guides users to interpret the scenarios in a benign or positive direction without having to actively and effortfully generate an alternative thought, as is CBT requires. For instance, here is an example CBM-I training scenario: ‘ You complete your final exam. Your teacher provides some feedback to you. You sense that the feedback is…h_lpf_l ’. In this example, the first sentence, ‘ You complete your final exam’ sets the scene; the second sentence, ‘ Your teacher provides some feedback to you’ creates the ambiguity designed to elicit a nature interpretation that is common to the individual—for people with depression, the interpretation is often negative (e.g., the teacher’s feedback is critical and to mean I’m useless; Rude et al., 2003); the final sentence, ‘ You sense that the feedback is…h_lpf_l’ presents the fragmented final word to deliver the word task. The client fills in the first missing letter of the fragmented final word to reveal it (i.e., helpful ), which resolves the ambiguity of the scenario in a more positive manner (see Figure 1). In a typical CBM-I training session, people are repeatedly exposed to a series of independent excerpts of this similar format. Researchers have demonstrated that CBM-I is effective in reducing negative interpretation bias in depression (Hirsch et al., 2018), as well as for various other psychological issues (e.g., paranoia, Yiend et al., 2023; social anxiety, Liu et al., 2017). Furthermore, in some studies, it has been shown that CBM-I is as effective as computerized-CBT in treating depression and anxiety (Bowler et al., 2012). Recent systematic review and meta-analysis studies have revealed that CBM-I had a moderate therapeutic effect on depression (Li et al., 2023; Martinelli et al., 2022) and other various mental health issues such as anxiety, eating disorders, and substance use (Martinelli et al., 2022; cf. Blackwell et al., 2015; Carlbring et al., 2012). In an earlier meta-analysis, however, researchers have shown that CBM-I produced only a small effect on anxiety and depression symptoms, and this effect was reliable only when symptoms were assessed after exposure to a stressor. When analyzed separately, CBM-I significantly reduced anxiety but not symptoms of depression (Hallion & Ruscio, 2011). Several reasons may account for the mix findings on the effects of CBM-I. One key feature of CBM-I training scenarios is content specificity—how closely the scenarios match everyday experiences that individuals with a specific mental health condition typically experience (e.g., depression). Researchers have demonstrated that using content specific scenarios have greater therapeutic power than using generalized content (Mackintosh et al., 2013; Savulich et al., 2017; Vancleef & Peters, 2008). We often see the same importance of using content specific materials in CBT (Beck, 1976). To develop content-specific CBM-I materials for depression, development phase of such materials typically involves an iterative process of working with relevant stakeholders, including experts by experience (i.e., people with living/lived experience of depression), clinicians, and researchers (Hsu et al., 2023). This can be costly and time exhaustive. To add to this complexity, CBM-I materials are presented in a three-sentence standardized format. CBM-I also includes second-person, gender-neutral pronouns (e.g., you) to enhance modification effects by creating an immersion effect to the scenarios (Steinman et al., 2021). Finally, to reflect the scheduled delivery of traditional psychological interventions, researchers have aimed for 6-12 weekly sessions consisting of 40 training items per session, yielding a total of 240-480 scenarios (Yiend et al., 2023). Given the specificity and number of training items required for effective treatment results, the development phase of CBM-I can become costly and resource-intensive, which may hinder its clinical availability and dissemination. In one study, Hsu et al. (2023) outlined the complex development process of CBM-I for treatment of paranoia. In that study, the material development process spanned over 12 months, involving people with first-hand experience of paranoia, clinicians, and a team of researchers. Materials were created and refined in an iterative manner, yielding a total of 240 scenarios. In another study, a similar process of creating and refining training materials for CBM-I reflected similar complexity (Hsu & Akuhatahuntington, 2024). The Potential Role of AI in CBM-I Material Development The recent advancement of generative AI chatbots may address cost issues involved in the materials development phase of CBM-I by making this process more efficient, quicker, and cost-effective (Alanezi, 2024; Montazeri et al., 2024), thereby improving the clinical availability and dissemination of the treatment; this is achieved through automation. Generative AI, such as ChatGPT, may be utilized to automatically create large quantities of CBM-I training scenarios that reflect human experiences of depression. Researchers have demonstrated that generative AI can accurately respond to open questions in a psychological medicine exam (Lin et al., 2024) and is best at repetitive tasks, such as generating text summaries rather than in performing critical thinking tasks (Duong & Solomon, 2023; Jeyaraman et al., 2023). Generative AI chatbots utilizes machine learning to produce texts in nature human language by drawing from large human data sets, which would likely include humans with lived depressive experiences, to pre-train transformer neural networks. Specialized algorithms involving reinforcement learning and reward models predict and produce contextually relevant content (Huh, 2023). Given that AI chatbots are trained using existing human data sets, and that they produce outputs in nature human language, AI chatbots could be used to create CBM-I training scenarios that reflect human experiences of depression. In other words, we could prompt AI chatbots to produce content specific scenarios reflecting common experiences of individuals experiencing depression from existing human data. This idea of using generative AI chatbots in CBM-I materials development is relatively novel. In only one study, researchers have investigated the use of a previous version of ChatGPT (v3.5) in relation to rating interpretation bias assessment materials. In that study, Lin et al. (2023) compared ChatGPT’s ratings to human ratings of scenarios used in a well-known interpretation bias measure known as the Similarity Rating Task. Scenarios, along with relevant biased interpretations, were developed by five human participants and then rated by another nine human participants and nine ChatGPT sessions. Ratings were benchmarked against a set of predefined criteria based on the degree of bias and readability of scenarios, on a 7-point scale to examine the equivalence of ratings between the two types of raters. The aim was to investigate whether ChatGPT rated the scenarios to the same extent as do humans. Despite showing a clear trend that both types of raters provided consistent ratings of the items in accordance with the direction of the biased interpretation, Lin et al.’s study was underpowered, making it difficult to interpret their results. The Present Study The overarching goal of the present study is to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials. By doing so, the overall cost of creating CBM-I materials may be significantly reduced and efficiency increased, thereby promoting its implementation on a broader scale and making it more accessible to both clinicians and clients as a treatment option for depression. We examined whether Copilot (https://copilot.microsoft.com), Microsoft’s state-of-the-art generative AI chatbot powered by the latest version of ChatGPT (v4.0), is capable of creating scenarios that capture situations and interpretation bias commonly associated with human experiences of depression. There is yet to be a study that has investigated the use of Copilot in creating CBM-I training materials; hence, this study is exploratory with no specific hypothesis. The study has the following objective: to compare participants’ ratings of human-generated and AI-generated raw and CBM-I materials, which included scenarios with a negative and positive interpretation. Method Procedure This study received institutional ethical approval (22/140). The present study included two stages which follows common development process of CBM-I training materials outlined in Hsu et al. ( 2023 ). Stage I involved the development of raw materials; Stage II involved adapting raw materials into CBM-I format (see Fig. 2 for a schematic outline of this process). This process involved using both human participants and Copilot to create and evaluate raw and CBM-I materials. Materials were validated on a set of predefined criteria: • relevance of scenarios in reflecting experiences common to individuals with depression • severity of the interpretation of scenarios • readability of scenarios (for CBM-I materials only) These criteria were selected to assess whether scenarios are content specific to individuals with depression ( relevance ) and that the interpretations of each scenario effectively capture a core cognitive feature of depression—distorted negative thinking patterns ( severity ). Training materials were assessed for the level of clarity and comprehensiveness of scenarios ( readability ). Each rating score is benchmarked against ratings of human-generated scenarios, which people with living/lived experience of depression creates. A deviation from human ratings in any direction on the rating scale would attend to our research question—that AI-generated materials are not equivalent to human-generated materials. By following the materials development process of CBM-I outlined in Hsu et al., we can effectively evaluate the potential of generative AI in replicating the development of CBM-I content. Participants Participants in both stages of the study were adults with self-report of living/lived experience of depression (advertised as “adults with first-hand experience of depression”). As this is a non-clinical study, we did not specifically recruit participants with a formal diagnosis of depression nor did we include screening measures of depression; instead, participants were recruited from the community in New Zealand through social media pages, a local job seeking website, and flyer advertisements posted in various public domains, including hospitals, university campus, and community clinics. The number of participants that entered the study varied across the different stages of the study. Prior to rating, all participants completed the Hospital Anxiety and Depression Scale (HADS; Zigmond & Snaith, 1983 ) to measure their current level of depression and anxiety for the purpose of providing a description of participant mood at the time of the study. The HADS is a 14-item self-report questionnaire designed to measure the severity of depression and anxiety using a 4-point Likert scale, with final scores ranging from 0–21 for each scale and grouped according to the degree of severity (normal = 0–7; borderline = 8–10; abnormal = 11+). See Table 1 for a description of the raters. Table 1 Raters’ Characteristics Rater: Raw Materials (n = 30) Rater: CBM-I Materials (n = 30) Depression Duration [Mean Years (SD)] 8.80 ( 7.90 ) 11 ( 8.78 ) HADS Score Depression [Mean score ( SD) ] 12.70 ( 3.26 ) 10.13 ( 5.16 ) Anxiety [Mean score ( SD) ] 8.03 ( 3.74 ) 13.43 ( 4.16 ) Gender (female:male:other) 23:3:4 26:3:1 Ethnicity (NZ European:Māori:other) 18:5:11 20:4:13 Age (Years) Mean ( SD ) 26.47 ( 9.00 ) 30.38 ( 8.66 ) Note. The ethnicity ratio does not equate to N = 30 because some raters identified with one or more ethnicity group. Stage I: Raw Materials Raw Materials Creation As a part of another study on CBM-I targeting depression, experts with living/lived experience of depression (N = 13) independently created a set of 130 raw items to describe their common everyday experiences (see Supplementary material, Table 1 ). Instructions given to the participants were “ write down 10 situations that could provoke depressive/negative interpretation; also provide a neutral/positive interpretation of that situation. The situation could be based on personal experience, someone else’s experience, or a situation that you can imagine experiencing.” Here is an example scenario: A friend hasn’t replied to a text , and its interpretations: They’re ignoring me because they don’t like me (negative); they might be busy and will reply when they have time (neutral/positive). The negative interpretation is created only to illustrate the ambiguity of the scenario and support its face validity—it is not used in the CBM-I training itself. The research team quasi-randomly selected 50 out of 130 raw items to be included in the present study, with 3 to 4 scenarios coming from each participant; we adopted a simple random sampling method using a random number generator to select scenarios that were created by each participant. This approach of selecting items maximized the variety of items to reflect depressive experiences that form part of the CBM-I intervention. Copilot was prompted to generate another set of 50 items to reflect human experiences of depression (see Supplementary Material, Table 2 ). Copilot has three conversation style options: ‘Precise, ‘Creative’, and ‘Balanced’. The ‘Precise’ conversation style was selected to optimize the precision of scenarios to human experience of depression. The instructions given to Copilot were similar to the instructions provided to human participants: “ Create different situations that someone with depression would commonly encounter and provide a depressive/negative interpretation and a neutral/positive interpretation of the situation”. Ten emotionally neutral items were created by the research team and included in the rating phase to serve as a manipulation check to ensure that raters are responding consistently and as expected. Manipulation check items were trivial statements designed to be unrelated to any emotional content (e.g., Phones, first appearing in 1849, have evolved over time ) presented with two sentences to complete the statement (e.g., You’re discussing the history of communication devices; You’re exploring the evolution of technology ). Since these items are emotionally neutral, the expectation is that raters should give similar ratings to both sentences. Hence, any differences in ratings of check items may suggest the presence of confounding variables in influencing the overall ratings. Raw Item Rating When developing these materials, it is common practice to refine content by working with living/lived experience experts to rate training content based on a set of pre-defined criteria. During the rating phase, all raw items were interleaved with 10 manipulation check items, yielding a total of 110 items. Scenarios were presented first, followed by two interpretations, one negative and one positive, that provided a different explanation to each scenario. Here is an example set that includes the situation along with the two interpretations: Situation : A friend hasn’t replied to a text. Here are the two possible interpretations of this situation: Interpretation 1 : They’re ignoring me because they don’t like me. Interpretation 2 : They might be busy and will reply when they have time. The interpretations were presented in random order, and the order was blinded to the raters. The raters (N = 30) rated all items independently and remotely on Qualtrics—an online survey platform. A priori power calculation using G*Power3.1 (Faul et al., 2009) revealed > 80% power at α = .05 using an equivalence bound of d = 0.718 (Lakens, 2017). The equivalence bound was derived from earlier work that compared ChatGPT ratings to human ratings of scenarios used in the Similarity Rating Task (see Lin et al., 2023 ). Items were rated against the pre-defined rating criteria using a 7-point scale: ‘relevance’ and ‘severity’. A higher rating on the ‘relevance’ scale is defined as the item being more representative of the experiences of individuals with depression (i.e., 1 = the least relevant to 7 = the most relevant). A higher rating on the ‘severity’ scale is defined as the item being more representative of the experiences of individuals with more severe depression (i.e., 1 = the least negative to 7 = the most negative). Each item includes the situation and one of the two interpretations. Stage II: CBM-I Items CBM-I Item Adaptation The human-generated raw items were adapted into CBM-I standard format by the research team; The AI-generated raw items were adapted into CBM-I format by Copilot (see Supplementary Material, Table 3; Table 4; Table 5). For the manipulation check items, half were adapted into CBM-I format by researchers, and the other half were adapted by Copilot. Here is an example of an adapted CBM-I scenario: You send a text message to a friend. They have not replied yet. You think they are…. and its interpretations: Ignoring you (negative); Busy (neutral/positive). Again, the negative interpretation is created only to illustrate the ambiguity of the scenario and support its face validity. The ‘Precise’ conversation style of Copilot was selected. Instructions provided to Copilot were “ Cognitive Bias Modification (CBM) format is 3 sentences long with the final word resolving the scenario in a negative manner or a positive manner. In CBM format, the interpretations are preferably 1 word but can be 2 or 3 if needed to make the sentence more readable.” An example of CBM-I format was also provided: “ Situation: You make a mistake at work. You receive some feedback about it. You think you are…Negative Interpretation: incompetent; Neutral/Positive Interpretation: improving” along and with additional instructions to support this process (e.g., “ CBM format keeps all three sentences identical except the last word is changed. make the CBM format like this:” and the situation should be in three sentences” ). CBM-I Item Rating Similar to the rating process for raw items, all CBM-I items were interleaved with 10 manipulation check items, yielding a total of 110 items. Scenarios were presented first, followed by two interpretations, one negative and one positive, that completed each scenario. For example: Situation : You send a text message to a friend. They have not replied yet. You think they are… Here are the two possible interpretations of this situation: Interpretation 1 : Ignoring you Interpretation 2 : Busy The interpretations were presented in random order, and the order was blinded to the raters. Another group of 30 raters were recruited from the same sources using the same method as the raters from the raw items rating phase (see Table 1 for raters’ characteristics). Raters rated all items independently and remotely on Qualtrics. Items were rated against the pre-defined rating criteria using a 7-point scale: ‘relevance’ and ‘severity’. An additional ‘ readability ’ criterion was included. A higher rating on this criterion reflected more clear and comprehensible scenarios. Results Given the exploratory nature of the present study, we conducted equivalence testing using the Two One-Sided Test (TOST) for paired data (Schuirmann, 1987 ) and null hypothesis significance testing (NHST) to analyze content ratings. Equivalence testing included an equivalence bound of Cohen’s d = -0.718 to d = 0.718 with a 90% confidence interval (see Lin et al., 2023 ). Research data are available here: https://tinyurl.com/GENAICBM . For analyses, we reversed the scale so that 1 = the least relevant to depression and most positive to 7 = the most relevant to depression and negative . This is to allow for more intuitive interpretation of rating scores. Stage I: Raw Items As shown in Fig. 3 left panel, results of TOST for paired data showed that human-generated and Copilot-generated raw items were statistically non-equivalent across all rating criteria (see Table 2 a for statistics). NHST using paired t-tests revealed convergent results in that ratings of human-generated and Copilot-generated raw items were significantly different across all rating criteria, with a difference in mean ratings falling between 0.1–0.48 (see Table 2 a). Statistically equivalent ratings were found for manipulation check items, suggesting that confounding effects of the order of presentation of scenarios with a positive and negative interpretation were likely controlled. Despite these findings, the direction of participants’ ratings aligned with the emotional valence of the scenarios. Specifically, both AI- and human-generated negative scenarios were rated above the midpoint on the criteria ‘relevance’ (M AI-generated = 5.93; M human-generated = 5.63) and ‘severity’ (M AI-generated = 6.06; M human-generated = 5.89), leaning toward the higher end. This suggests consistency in how both sources generated excerpts that conveyed negative content. Similarly, positive scenarios were rated closer to the lower end of the scale (relevance: M AI-generated = 2.91; M human-generated = 3.39 and severity: M AI-generated = 2.45; M human-generated = 2.76), reflecting their intended positive valence. Furthermore, ratings of manipulation check items clustered around the midpoint (relevance: M AI-generated = 4.61; M human-generated = 4.54 and severity: M AI-generated = 3.63; M human-generated = 3.56),further supporting the overall consistency in rating directionality. A careful examination of the ratings revealed that participants consistently rated the AI-generated items as more negative on the severity dimension and more relevant to common experiences of individuals with depression. A similar pattern of directionality was observed for positively valence items, with raters providing lower ratings on the AI-generated items. These results suggest that participants tended to rate AI-generated items were higher when they depicted negative content and lower when they depicted positive content, indicating a consistent directional trend in their evaluations. Table 2 a Results and Statistical Tests of Raw Items [Mean (SD), 90% confidence interval CI] Rating Criteria AI-Generated Human-Generated Equivalence Test Difference Test Relevance Negative Scenarios 5.93 (0.77) 5.63 (0.70) t(29) = 0.98, p = .83 , CI -0.41 to -0.2 t(29) = -4.91, p < .001, CI 0.20 to 0.41 Positive Scenarios 2.91 (1.02) 3.39 (0.97) t(29) = -3.15, p = .10 , CI 0.37 to 0.6 t(29) = 7.08, p < .001, CI -0.6 to -0.37 Manipulation check items 4.61 (0.87) 4.54 (0.84) t(29) = -3.10, p = .002 , CI -0.2 to 0.07 t(29) = -0.83, p = .41, CI-0.07 to -0.20 Severity : Negative interpretation 6.06 (0.78) 5.89 (0.73) t(29) = -0.37, p = .36 , CI -0.24 to -0.09 t(29) = -3.56, p = .001, CI 0.09 to -0.24 Positive interpretation 2.45 (0.90) 2.76 (0.70) t(29) = -0.92, p = .82 , CI 0.2 to 0.41 t(29) = 4.85, p < .001, CI -0.41 to -0.2 Manipulation check items 3.63 (0.52) 3.56 (0.53) t(29) = -2.71, p = .006 , CI -0.18 to 0.03 t(29) = 0.23, p = .23, CI -0.03 to 0.18 Stage II: CBM-I Items Following the adaptation of raw items into CBM-I format, we examined both the equivalence and differences in mean ratings of human-adapted and Copilot-adapted CBM-I items. The same statistical analyses used in comparing raw items were performed for this comparison (see Fig. 3 , right panel). With the exception of severity ratings, comparing human-adapted and Copilot-adapted CBM-I items resulted in statistically non-equivalent (and statistically different) ratings, with a difference in mean ratings falling between 0.03–0.41 (see Table 2 b). Similar to the ratings of the raw items, participants’ ratings for the relevance criterion—but not severity—generally aligned with the emotional content of both the negative and positive CBM-I scenarios and the manipulation check items, indicating consistent directionality in how emotional content was perceived. Participants also consistently rated the AI-generated CBM-I items as more relevant to common experiences of individuals with depression (M AI-generated = 5.93; M human-generated = 5.51). A similar pattern of directionality was observed for positively valence items, with raters providing lower ratings on the AI-generated CBM-I items (M AI-generated = 3.45; M human-generated = 3.68). Finally, statistically equivalent ratings were observed for manipulation check items only on the relevance dimension only—not severity. This result suggests that the order in which positive and negative interpretations were presented may have influenced participants’ severity ratings, potentially explaining the inconsistencies in ratings we found between the relevance and severity scales. Taken together, mean ratings of both Copilot-generated and human-generated raw items and adapted CBM-I items were both statistically non-equivalent and different, with the exception of the severity ratings on CBM-I items. The overall direction of ratings between AI-and human-generated scenarios were the same and consistent with the scenarios’ emotional content. Table 2 b Results and Statistical Tests of CBM-I Items [Mean (SD), 90% confidence interval CI] Rating Criteria AI-Generated Human-Generated Equivalence Test Difference Test Relevance: Negative Scenarios 5.93 (0.77) 5.51 (0.84) t(29) = 2.79, p = .10 , CI -0.52 to -0.31 t(29) = -6.72, p < .001, CI 0.31 to 0.52 Positive Scenarios 3.45 (1.08) 3.68 (1.1) t(29) = -0.29, p = .61 , CI 0.13 to 0.31 t(29) = 4.22, p < .001, CI -0.31 to -0.13 Manipulation check items 3.94 (0.68) 4.04 (0.60) t(29) = 2.96, p = .003 , CI -0.08 to 0.28 t(29) = 0.97, p = .34, CI -0.28 to 0.08 Severity : Negative interpretation 5.62 (0.93) 5.65 (0.95) t(29) = 3.18, p = .002 , CI -0.05 to 0.12 t(29) = 0.75, p = .46, CI -0.12 to 0.045 Positive interpretation 3.10 (0.77) 3.02 (0.83) t(29) = -2.42, p = .01 , CI -0.17 to 0.01 t(29) = -1.51, p = .14, CI -0.01 to 0.17 Manipulation check items 4.18 (0.62) 3.92 (0.44) t(29) = − .81, p = .21 , CI − 0.41 to -0.12 t(29) = -3.13, p = .004, CI 0.12 to 0.41 Readability : 5.91 (1.47) 5.87 (1.47) t(29) = 1.49, p = .07 , CI -0.08 to -0.01 t(29) = 2.44, p = .02, CI 0.01 to 0.08 Discussion Cognitive bias modification-interpretation (CBM-I) is a digital therapeutic that utilizes ambiguous text-based scenarios to capture and modify damaging interpretation bias known to contribute to depression. An integral part of CBM-I’s success in modifying bias is the quality of training scenarios. The development process of these key materials is often onerous and costly, involving an iterative process of extensive input from multiple stakeholders to create and refine training content. The aim of the study was to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials, as per pre-defined rating criteria. We followed a typical development procedure, as outlined in Hsu et al. ( 2023 ), to create CBM-I items; first, developing raw items and then adapting the raw items into CBM-I format. Items were generated by AI (CoPilot) and human experts (i.e., people with living/lived experience of depression). We compared experts’ ratings of AI-generated and human-generated items. With a larger sample size and using a more powerful AI platform compared to ChatGPT3.5 (Wang et al., 2023 ), our results offer a new perspective on Lin et al.’s ( 2023 ) underpowered study where the authors used GPT3.5 to rate scenarios presented in the Similarity Rating Task, a bias assessment that is similar to the format of CBM-I. In our study, we found statistically non-equivalent and significant differences in mean ratings between human- and AI-generated raw and CBM-I items. There is a paucity of studies on generative AI in creating mental health scenarios, and the closest comparison we could draw from are studies on ChatGPT’s capabilities in generating clinical vignettes for medical education. In this area of research, concerns have been raised regarding ChatGPT’s accuracy and reliability in producing clinical scenarios that accurately depict common medical presentations (Alam et al., 2023 ), possibly due to limitations in its training data (Sahu et al., 2023 ). In a scoping review, Xu et al. ( 2024 ) found that, in 42.5% of the reviewed articles, researchers have discussed that ChatGPT may present incorrect or inconsistent medical information and concepts in AI-generated clinical vignettes. Despite our data indicating statistical non-equivalence in the overall ratings of human-generated and AI-generated materials, several important caveats must be considered when interpreting these findings. First, our study found no statistical differences in participants’ ratings of severity of CBM-I items; instead, statistical equivalence was established. One possible explanation for this result is that the manipulation check items failed to show equivalence, suggesting that confounding variables may have influenced how participants rated the items. Next, caution is needed when generalizing results from studies that used AI-generated clinical vignettes, as the vignettes may be fundamentally different from people’s experience of depression and serve different purposes. Furthermore, many of these studies adopted a weaker study design, such as comparing participants’ ratings of excerpts rather than using other more rigorous methods like randomized control trials (cf. Coşkun et al., 2024 ). Similar to our study, ratings are often based on a limited set of criteria used to evaluate the similarity between AI-generated and human-generated materials. However, other factors—such as accuracy, readability, and clinical outcomes—as well as broader considerations, including safety and societal implications, are also crucial in assessing the effectiveness, reliability, and quality of AI-generated materials. These dimensions should be included in future avenues of study. Finally, when attending to our research question on whether AI-generated CBM-I training materials are equivalent to human-generated materials, it is important to consider that deviations from ratings of human-generated materials—the benchmark in our study—do not necessarily undermine the usefulness of AI-generated content (Lin et al., 2023 ). This is particularly relevant when AI- and human-generated excerpts show similar rating patterns that align with the intended emotional content. In our study, both sources generated excerpts that were either consistent with negative or positive scenarios. In most cases, participants also consistently rated AI-generated items higher when they depicted negative content and lower when they depicted positive content, indicating a consistent directional trend in their evaluations. Moreover, it is important to examine results beyond statistical testing and consider the definition of equivalence in a clinical setting. For instance, by considering actual differences in mean ratings of materials, our data showed that the mean ratings fell within one unit/point on a 7-point rating scale between 0.03–0.41, suggesting that the differences were minimal and may have limited impact clinically. More specifically, our data on the relevance criterion showed that mean ratings for both AI-generated items (M = 5.93) and human-generated items (M = 5.51) were above the mid-point for negative interpretations; ratings of positive interpretation were below the mid-point (M AI−generated = 3.45; M human−generated = 3.68). In a medical education study, Benoit ( 2023 ) demonstrated that, despite ChatGPT creating clinical vignettes that included additional disease symptoms that was beyond what it was instructed to provide, the extraneous information depicted an accurate clinical symptom. Taken together, despite the statistically non-equivalence of participants’ ratings of human-generated and AI-generated materials, when interpreting the data, researchers should consider other confounding variables and evaluate what ‘true equivalence’ really means. Future directions of research should consider using more rigorous study designs—a randomized control trial—and aim to examine the implementation of AI-generated items for CBM-I against human-generated items. Furthermore, carefully analyzing the excerpts and comparing them qualitatively may provide more rich understanding of AI-generated items. Finally, given the rapid advancement of generative AI technology, more recent or future versions of AI chatbots may have enhanced capabilities of generating scenarios that closely align with human experiences of depression. The benefits of using generative AI chatbots to create CBM-I training materials are beyond that of automation. Generative AI chatbots uses specialized algorithms to seek out patterns from human prompts and uses reinforcement learning and reward models to provide a response. The implication of this would mean individualised CBM-I training materials tailored to individual service users (Ahmad et al., 2022 ), which may improve the content specificity of materials and promote better bias modification effects (Mackintosh et al., 2013 ; Savulich et al., 2017 ; Vancleef & Peters, 2008 ). What this may look like practically warrants further investigation and encouraged in future avenue of studies. It could mean that the chatbot is trained using individualized data to create similar descriptions of the individual’s experiences and biases. There are some limitations to the present study that should be considered in future studies. One major issue is the representation of the scenarios. That is, the validity of generalizing experiences from 13 individuals with depression may be low. In future studies, it is encouraged that researchers gather everyday experiences from a larger and diverse population, considering age, gender, ethnicity, and years of depression. An additional limitation is that only simple prompting was used to instruct Copilot in creating scenarios. Prompts in AI are pivotal for optimizing outputs (Bozkurt, 2024 ) and several strategies have been researched and developed (Liu et al., 2023 ). It is important to test different prompt strategies to determine the optimal prompt for creating CBM-I training items. Finally, although the use of AI-generated scenarios may reduce cost associated with developing CBM-I materials, one should consider other potential cost of using AI, such as energy and environmental cost. A cost-benefit analysis may be warranted in future studies. Conclusion In conclusion, this study provides initial insights into the comparison between human-generated and AI-generated CBM-I materials. While the findings revealed AI- and human-generated scenarios are statistically non-equivalent, as per pre-defined rating criteria, it is important to recognize the potential of AI-generated materials on a more practical level. On this basis, our findings should be viewed as part of a broader and ongoing research in AI’s capabilities of generating excerpts for clinical use. Future research with a larger sample size and clinical validation would be necessary to better understand the effectiveness of AI in this context. Using generative AI in developing treatment materials for CBM-I training for depression could significantly reduce overall costs and increase the efficiency of the materials development process. This, in turn, could facilitate the broader dissemination and implementation of an effective intervention for depression. References Ahmad, R., Siemon, D., Gnewuch, U., & Robra-Bissantz, S. (2022). Designing personality-adaptive conversational agents for mental health care. Information Systems Frontiers, 24 (3), 923–43. https://doi.org/10.1007/s10796-022-10254-9 Alam, F., Lim, M. A., & Zulkipli, I. N. (2023). Integrating AI in medical education: embracing ethical usage and critical understanding. Frontiers in Medicine, 10, article 1279707. https://doi.org/10.3389/fmed.2023.1279707 Alanezi, F. (2024). Assessing the Effectiveness of ChatGPT in Delivering Mental Health Support: A Qualitative Study. Journal of Multidisciplinary Healthcare, 17, 461–471. https://doi.org/10.2147/JMDH.S447368 Blackwell, S. E., Browning, M., Mathews, A., Pictet, A., Welch, J., Davies, J., et al. (2015). Positive imagery-based cognitive bias modification as a web-based treatment tool for depressed adults: A randomized controlled trial. Clinical Psychological Science, 3 (1), 91–111. https://doi.org/10.1177/2167702614560746 Beck, A. T. (1976). Cognitive therapy and the emotional disorders. International Universities Press. Benoit, J. R. A. (2023). ChatGPT for clinical vignette generation, revision, and evaluation. medRxiv. https://doi.org/10.1101/2023.02.04.23285478 Bowler, J. O., Mackintosh, B., Dunn, B. D., Mathews, A., Dalgleish, T., & Hoppitt, L. (2012). A comparison of cognitive bias modification for interpretation and computerized cognitive behavior therapy. Journal of Consulting and Clinical Psychology, 80 (6), 1021–1033. https://doi.org/0022-006X/12/$12.00 Bozkurt (2024). Tell me your prompts and I will make them true: the alchemy of prompt engineering and generative AI. Open Praxis, 16 (2), 111–118. https://doi.org/10.55982/openpraxis.16.2.661 Carlbring, P., Apelstrand, M., Sehlin, H., Amir, N., Rousseau, A., Hofmann, S. G., et al. (2012). Internet-delivered attention bias modification training in individuals with social anxiety disorder: A double-blind randomized controlled trial. BMC Psychiatry, 12 , 66. https://doi.org/10.1186/1471-244X-12-66 Chen J., Short, M., & Kemps, E. (2020). Interpretation bias in social anxiety: A systematic review and meta-analysis. Journal of Affect Disorder, 276 (3) , 1119–1130. https://10.1016/j.jad.2020.07.121 Coskun, O., Kiyak, Y. S., & Budakoglu, I. I. (2024). ChatGPT to generate clinical vignettes for teaching and multiple-choice questions for assessment: A randomized controlled experiment. Medical Teacher, 13, 1–7. https://doi.org/10.1080/0142159X.2024.2327477 Duong, D. & Solomon, B. D. (2023). Analysis of large-language model vs human performance for genetics questions. European Journal of Human Genetics, 32 (4), 466–468. https://doi.org/10.1038/s41431-023-01396-8 Everaert, J., Podina, I. R., & Koster, E. H. W. (2017). A comprehensive meta-analysis of interpretation biases in depression. Clinical Psychology Review, 58, 33–48. https://doi.org/10.1016/j.cpr.2017.09.005 Garety, P. A., Kuipers, E., Fowler, D., Freeman, D., & Bebbington, P. E. (2001). A cognitive model of the positive symptoms of psychosis. Psychological Medicine , 31 (2), 189–195. https://doi.org/10.1017/s0033291701003312 Ha, S.W. & Kim, J. (2020). Designing a scalable, accessible, and effective mobile app based solution for common mental health problems. International Journal of Human-Computer Interaction, 36 (35), 1354–1367. https://doi.org/10.1080/10447318.2020.1750792 Hallion, L. S., & Ruscio, A. M. (2011). A meta-analysis of the effect of cognitive bias modification on anxiety and depression. Psychological Bulletin, 137 (6), 940–958. https://doi.org/10.1037/a0024355 Hirsch, C. R., Krahé, C., Whyte, J., Loizou, S., Bridge, L., Norton, S., & Mathews, A. (2018). Interpretation training to target repetitive negative thinking in generalized anxiety disorder and depression. Journal of Consulting and Clinical Psychology, 86 (12), 1017–1030. https://doi.org/10.1037/ccp0000310 Hsu, C. W. & Akuhatahuntington, Z. (2024). Implicit bias training for New Zealand medical students using cognitive bias modification: An outline of material development. New Zealand Journal of Psychology , 52 (3), 30–43. Hsu, C. W., Stahl, D., Mouchlianitis, E., Peters, E., Vamvakas, G., Keppens, J., Watson, M., Schmidt, N., Jacobsen, P., McGuire, P., Shergill, S., Kabir, T., Hirani, T., Yang, Z. Y., & Yiend, J. (2023). User-Centered Development of STOP (Successful Treatment for Paranoia): Material Development and Usability Testing for a Digital Therapeutic for Paranoia. JMIR Human Factors, 10, e45453. https://doi.org/10.2196/45453 Huh, S. (2023). Are ChatGPT’s knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination? A descriptive study. Journal of Educational Evaluation of Health Professions, 20 , 1. https://doi.org/10.3352/jee­hp.2023.20.1 Jeyaraman, M., Ramasubramanian, S., Balaji, S., Jeyaraman, N., Nallakumarasamy, A., & Sharma, S. (2023). ChatGPT in action: Harnessing artificial intelligence potential and addressing ethical challenges in medicine, education, and scientific research. World Journal of Methodology, 13 (4), 170–178. https://doi.org/10.5662/wjm.v13.i4.170 Lakens, D. (2018). Equivalence Testing for Psychological Research: A Tutorial. Advances in Methods and Practices in Psychological Science, 1 (2), 259–269. https://doi.org/10.1177/2515245918770963 Li J. W., Ma, H., Yang, H., Yu, H., & Zhang, N. (2022). Cognitive bias modification for adult’s depression: A systematic review and meta-analysis. Frontiers in Psychology, 13, 968638. https://10.3389/fpsyg.2022.968638 Lin, C., C., Akuhata-Huntington, Z., & Hsu, C. W. (2023). Comparing ChatGPT’s ability to rate the degree of stereotypes and the consistency of stereotype attribution with those of medical students in New Zealand in developing a similarity rating test: a methodological study. Journal of Educational Evaluation of Health Professions, 20 , 17. https://10.3352/jeehp.2023.20.17 Lin, C. C., du Plooy, K., Gray, A., Brown, D., Hobbs, L., Patterson, T., Tan, V., Fridberg, D., Hsu, C. W. (2024). The performance of ChatGPT on short-answer questions in a psychiatry examination: A pilot study. Taiwanese Journal of Psychiatry, 38 (2), 94–98. https://doi.org/10.4109/TPSY.TPSY_19_24 Liu, H., Li, X., Han, B, & Liu, X. (2017). Effects of cognitive bias modification on social anxiety: A meta-analysis. PLoS One, 12 (4): e0175107 Liu, P. F., Y, W. Z., Fu, J. L., Jiang, Z. B., Hayashi, H., & Neubig, G. (2023). Pre-train, prompt, and predict. A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55 (9), 1–35. https://doi.org/10.1145/3560815 Mackintosh, B., Mathews, A., Eckstein, D., & Hoppitt, L. (2013). Specificity effects in the modification of interpretation bias and stress reactivity. Journal of Experimental Psychopathology, 4 (2) , 133–147 . https://doi.org/10.5127/jep.025711 Martinelli, A., Grüll, J., & Baum, C. (2022). Attention and interpretation cognitive bias change: A systematic review and meta-analysis of bias modification paradigms. Behaviour Research and Therapy, 157, 104180. https://doi.org/10.1016/j.brat.2022 Mathews, A. & Mackintosh B. (2000). Induced emotional interpretation bias and anxiety. Journal of Abnormal Psychology, 109 (4), 602–615. https://doi.org/10.1037/0021-843X.109.4.602 Montazeri, M., Galavi, Z., & Ahmadian, L. (2024). What are the applications of ChatGPT in healthcare: Gain or loss? Health Science Reports, 7 (2), e1878. https://10.1002/hsr2.1878 Rude, S., Valdez, C.R., Odom, S., & Ebrahimi, A. (2003). Negative cognitive biases predict subsequent depression. Cognitive Therapy and Research, 27 (4), 415–429. https://doi.org/10.1023/A:1025472413805 Sahu, P. K., Benjamin, L. A., Singh, A. G., & Williams-Persad, A. (2023). ChatGPT in research and health professions education: challenges, opportunities, and future directions. Postgraduate Medical Journal, 100 (1179), 50–55. https://doi.org/10.1093/postmj/qgad090 Savulich, G., Freeman, D., Shergill, S., & Yiend, J. (2015). Interpretation biases in paranoia. Behavior Therapy, 46 (1), 110–124. https://doi.org/10.1016/j.beth.2014.08.002 Savulich, G., Shergill, S. S., & Yiend, J. (2017). Interpretation biases in clinical paranoia. Clinical Psychological Science , 5 (6), 985–1000. https://doi.org/10.1177/2167702617718180 Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Pharmacodynamics, 15 (6), 657–680. https://doi.org/10.1007/BF01068419 Steinman, S. A., Namaky, N., Toton, S. L., Meissel, E. E. E., John, A. T. S., Pham, N. H., Werntz, A., Valladares, T. L., Gorlin, E. I., Arbus, S., Beltzer, M., Soroka, A., & Teachman, B. A. (2021). Which variations of a brief cognitive bias modification session for interpretations lead to the strongest effects? Cognitive Therapy and Research, 45 (2), 367–382. https://doi.org/10.1007/s10608-020-10168-3 Vancleef, L. M. G. & Peters, M. L. (2008). Examining content specificity of negative interpretation biases with the Body Sensations Interpretation Questionnaire (BSIQ). Journal of Anxiety Disorder, 22 (3), 401–415. https://doi.org/10.1016/j.janxdis.2007.05.006. Wang, G., Gao, K., Liu, Q. Y., Wu, Y. X., Zhang, K. J., Zhou, W., & Gui, C. B. (2023). Potential and Limitations of ChatGPT 3.5 and 4.0 as a Source of COVID-19 Information: Comprehensive Comparative Analysis of Generative and Authoritative Information. Journal of Medical Internet Research, 25, e49771. https://doi.org/10.2196/49771 Xu, X. J., Chen, Y. X., & Miao, J. (2024). Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review. Journal of Educational Evaluation for Health Professions, 21(6). https://doi.org/10.3352/jeehp.2024.21.6 Yiend, J., Lam, C. L. M., Schmidt, N., Crane, B., Heslin, M., Kabir, T., McGuire, P., Meek, C., Mouchlianitis, E., Peters, E., Stahl, D., Trotta, A., & Shergill, S. (2023). Cognitive bias modification for paranoia (CBM-pa): a randomised controlled feasibility study in patients with distressing paranoid beliefs. Psychological Medicine, 53 (10), 4614–4626. https://doi.org/10.1017/S0033291722001520 Yiend, J., Lee, J. S., Tekes, S., Atkins, L., Mathews, A., Vrinten, M., Ferragamo, C., & Shergill, S. (2013). Modifying Interpretation in a Clinically Depressed Sample Using ‘Cognitive Bias Modification-Errors’: A Double Blind Randomised Controlled Trial. Cognitive Therapy and Research, 38 (2), 146–159. https://10.1007/s10608-013-9571-y Zigmond, A. S., & Snaith, R. P. (1983). The hospital anxiety and depression scale. Acta Psychiatrica Scandinavica, 67(6), 361–370. https://doi.org/10.1111/j.1600-0447.1983.tb09716.x Supplementary Material Supplementary Tables 1-5 are not available with this version. Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7530420","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":509940997,"identity":"5fbc032d-59d6-439b-8a97-ba422e308e22","order_by":0,"name":"Che-Wei Hsu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8klEQVRIiWNgGAWjYBACxgYwdYCBgb2BZC08BxiJ1QTTIpFApBbm9t5nn3kY7tgb3Hx8/MGPXzYM/O0H2B7+wOewnuPGs3kYniVuuJ2W2Njbl8YgcSaB3UACn5YZaczMPAyHEwxu5xg28PYcZmC4wcAmYYBPy/xnYC1Ah50xbPwL1CIP0pKA1xY2sBbGDTd4DJt5fhxmMABpOYDXL2nMjHMMDifOPJOWOFu2IY3H8Exim2QDHi2G7ceYGd5UHLbnO374wMc3f2zk5I4fPiaJL8QMgeYx8cB8y9jGwAOPXlxAHqQQYeYfvIpHwSgYBaNghAIAd1hN5aKvKq0AAAAASUVORK5CYII=","orcid":"","institution":"University of Otago","correspondingAuthor":true,"prefix":"","firstName":"Che-Wei","middleName":"","lastName":"Hsu","suffix":""},{"id":509940998,"identity":"83611207-5236-40ad-bcb2-f03601332ac2","order_by":1,"name":"Alex Robbins","email":"","orcid":"","institution":"University of Otago","correspondingAuthor":false,"prefix":"","firstName":"Alex","middleName":"","lastName":"Robbins","suffix":""},{"id":509940999,"identity":"852b0391-bfce-461c-a76b-15e968076e35","order_by":2,"name":"Tiana Cartwright","email":"","orcid":"","institution":"University of Otago","correspondingAuthor":false,"prefix":"","firstName":"Tiana","middleName":"","lastName":"Cartwright","suffix":""}],"badges":[],"createdAt":"2025-09-03 22:06:52","currentVersionCode":1,"declarations":{"humanSubjects":true,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":true,"humanSubjectConsent":true,"humanSubjectClinicalTrial":true,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7530420/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7530420/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":90893477,"identity":"e8ab5f66-cf00-48a5-a61a-6590485f8d0b","added_by":"auto","created_at":"2025-09-09 11:24:19","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":115455,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAn example of CBM-I training scenario targeting negative interpretation bias\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eNote. \u003c/em\u003eSurvey elements, Copyright Qualtrics, LLC. Used With Permission\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7530420/v1/38d5f96cedf4aa381a445f31.png"},{"id":90893476,"identity":"b4bfa4d0-efef-4158-8810-65bd52b29a15","added_by":"auto","created_at":"2025-09-09 11:24:19","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":26328,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSchematic of a typical co-development process of developing CBM-I training materials (adapted from Hsu et al., 2023).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7530420/v1/72e5489a81473c8455dacf6f.png"},{"id":90895207,"identity":"31ff9e2c-1860-4acc-995c-612169cc134a","added_by":"auto","created_at":"2025-09-09 11:32:46","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":19020,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eParticipants’ ratings of AI- and human-generated raw and CBM-I items\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7530420/v1/6f0dd15078bfde8c9abeba1a.png"},{"id":90897742,"identity":"7e84d8d9-9ab6-4658-8f2a-fde02c9300f3","added_by":"auto","created_at":"2025-09-09 11:48:48","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1014405,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7530420/v1/46c39e31-afbc-447d-91c2-c05ef884129a.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eGenerative artificial intelligence in mental health: A preliminary study on automating materials development for cognitive bias modification.\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAccording to cognitive models of psychological disorders, interpretation bias is thought to contribute to the formation and maintenance of various mental health concerns (Beck, 1976; Garety et al., 2001). Interpretation bias, in this context can be defined as the tendency to interpret ambiguous situations with a negative or threatening connotation. Across many studies, researchers have found compelling evidence of interpretation bias among people who experience various mental health issues such as depression (Everaert et al., 2017), anxiety (Chen et al., 2020), and paranoia (Savulich et al., 2015). More importantly from an intervention standpoint, it has been well-documented that interpretation bias can be modified to reduce mental health symptoms (Li et al., 2022; Yiend et al., 2013). Cognitive bias modification-Interpretation (CBM-I) is a class of evidence-based cognitive therapeutic designed to directly target damaging interpretation bias to improve symptomology. There are several features of CBM-I that may make it an appealing alternative to traditional psychological interventions such as cognitive behavioural therapy (CBT). First, CBM-I is self-administered via a mobile phone app or a web platform, which means the treatment is highly accessible, scalable, and discrete, making it easier to disseminate to a wider population. It is also designed to be a low-cost treatment (Ha \u0026amp; Kim, 2020), with cost referring to time, personnel, and financial expenses. Additionally, CBM-I can be simpler to implement than CBT and offers a more targeted approach to treatment, which may enhance its clinical application.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDespite these advantages over other therapeutics, developing training materials for CBM-I can be challenging and costly, thereby limiting its widespread utility and availability in clinical settings. The overarching goal of the present study is to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials. In this context, we define equivalence in a narrow sense, basing it on a set of pre-defined criteria designed to capture common experiences of individuals with depression.\u003c/p\u003e\n\u003cp\u003eCBM-I adopts findings from basic research of interpretation bias and its association with depression. More specifically, people with depression tend to interpret ambiguous everyday events in a more negative direction (Rude et al., 2003); using text-based excerpts of such events could successfully evoke and modify people\u0026rsquo;s negative interpretation bias to reduce depression (Li et al., 2022; Mathews \u0026amp; Mackintosh, 2000; Yiend et al., 2013). CBM-I employs a set of ambiguous scenarios designed to elicit multiple pathways of interpretation. By using a word task, CBM-I guides users to interpret the scenarios in a benign or positive direction without having to actively and effortfully generate an alternative thought, as is CBT requires.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor instance, here is an example CBM-I training scenario: \u0026lsquo;\u003cem\u003eYou complete your final exam. Your teacher provides some feedback to you. You sense that the feedback is\u0026hellip;h_lpf_l\u003c/em\u003e\u0026rsquo;. In this example, the first sentence, \u0026lsquo;\u003cem\u003eYou complete your final exam\u0026rsquo;\u003c/em\u003e sets the scene; the second sentence, \u0026lsquo;\u003cem\u003eYour teacher provides some feedback to you\u0026rsquo;\u003c/em\u003e creates the ambiguity designed to elicit a nature interpretation that is common to the individual\u0026mdash;for people with depression, the interpretation is often negative (e.g., the teacher\u0026rsquo;s feedback is critical and to mean I\u0026rsquo;m useless; Rude et al., 2003); the final sentence, \u0026lsquo;\u003cem\u003eYou sense that the feedback is\u0026hellip;h_lpf_l\u0026rsquo;\u003c/em\u003e presents the fragmented final word to deliver the word task. The client fills in the first missing letter of the fragmented final word to reveal it (i.e., \u003cem\u003ehelpful\u003c/em\u003e), which resolves the ambiguity of the scenario in a more positive manner (see Figure 1). In a typical CBM-I training session, people are repeatedly exposed to a series of independent excerpts of this similar format.\u003c/p\u003e\n\u003cp\u003eResearchers have demonstrated that CBM-I is effective in reducing negative interpretation bias in depression (Hirsch et al., 2018), as well as for various other psychological issues (e.g., paranoia, Yiend et al., 2023; social anxiety, Liu et al., 2017). Furthermore, in some studies, it has been shown that CBM-I is as effective as computerized-CBT in treating depression and anxiety (Bowler et al., 2012). Recent systematic review and meta-analysis studies have revealed that CBM-I had a moderate therapeutic effect on depression (Li et al., 2023; Martinelli et al., 2022) and other various mental health issues such as anxiety, eating disorders, and substance use (Martinelli et al., 2022; cf. Blackwell et al., 2015; Carlbring et al., 2012). In an earlier meta-analysis, however, researchers have shown that CBM-I produced only a small effect on anxiety and depression symptoms, and this effect was reliable only when symptoms were assessed after exposure to a stressor. When analyzed separately, CBM-I significantly reduced anxiety but not symptoms of depression (Hallion \u0026amp; Ruscio, 2011).\u003c/p\u003e\n\u003cp\u003eSeveral reasons may account for the mix findings on the effects of CBM-I. One key feature of CBM-I training scenarios is content specificity\u0026mdash;how closely the scenarios match everyday experiences that individuals with a specific mental health condition typically experience (e.g., depression). Researchers have demonstrated that using content specific scenarios have greater therapeutic power than using generalized content (Mackintosh et al., 2013; Savulich et al., 2017; Vancleef \u0026amp; Peters, 2008). We often see the same importance of using content specific materials in CBT (Beck, 1976). To develop content-specific CBM-I materials for depression, development phase of such materials typically involves an iterative process of working with relevant stakeholders, including experts by experience (i.e., people with living/lived experience of depression), clinicians, and researchers (Hsu et al., 2023). This can be costly and time exhaustive. To add to this complexity, CBM-I materials are presented in a three-sentence standardized format. CBM-I also includes second-person, gender-neutral pronouns (e.g., you) to enhance modification effects by creating an immersion effect to the scenarios (Steinman et al., 2021). Finally, to reflect the scheduled delivery of traditional psychological interventions, researchers have aimed for 6-12 weekly sessions consisting of 40 training items per session, yielding a total of 240-480 scenarios (Yiend et al., 2023).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGiven the specificity and number of training items required for effective treatment results, the development phase of CBM-I can become costly and resource-intensive, which may hinder its clinical availability and dissemination. In one study, Hsu et al. (2023) outlined the complex development process of CBM-I for treatment of paranoia. In that study, the material development process spanned over 12 months, involving people with first-hand experience of paranoia, clinicians, and a team of researchers. Materials were created and refined in an iterative manner, yielding a total of 240 scenarios. In another study, a similar process of creating and refining training materials for CBM-I reflected similar complexity (Hsu \u0026amp; Akuhatahuntington, 2024).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eThe Potential Role of AI in CBM-I Material Development\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe recent advancement of generative AI chatbots may address cost issues involved in the materials development phase of CBM-I by making this process more efficient, quicker, and cost-effective (Alanezi, 2024; Montazeri et al., 2024), thereby improving the clinical availability and dissemination of the treatment; this is achieved through automation. Generative AI, such as ChatGPT, may be utilized to automatically create large quantities of CBM-I training scenarios that reflect human experiences of depression. Researchers have demonstrated that generative AI can accurately respond to open questions in a psychological medicine exam (Lin et al., 2024) and is best at repetitive tasks, such as generating text summaries rather than in performing critical thinking tasks (Duong \u0026amp; Solomon, 2023; Jeyaraman et al., 2023). Generative AI chatbots utilizes machine learning to produce texts in nature human language by drawing from large human data sets, which would likely include humans with lived depressive experiences, to pre-train transformer neural networks. Specialized algorithms involving reinforcement learning and reward models predict and produce contextually relevant content (Huh, 2023). Given that AI chatbots are trained using existing human data sets, and that they produce outputs in nature human language, AI chatbots could be used to create CBM-I training scenarios that reflect human experiences of depression. In other words, we could prompt AI chatbots to produce content specific scenarios reflecting common experiences of individuals experiencing depression from existing human data.\u003c/p\u003e\n\u003cp\u003eThis idea of using generative AI chatbots in CBM-I materials development is relatively novel. In only one study, researchers have investigated the use of a previous version of ChatGPT (v3.5) in relation to rating interpretation bias assessment materials. In that study, Lin et al. (2023) compared ChatGPT\u0026rsquo;s ratings to human ratings of scenarios used in a well-known interpretation bias measure known as the Similarity Rating Task. Scenarios, along with relevant biased interpretations, were developed by five human participants and then rated by another nine human participants and nine ChatGPT sessions. Ratings were benchmarked against a set of predefined criteria based on the degree of bias and readability of scenarios, on a 7-point scale to examine the equivalence of ratings between the two types of raters. The aim was to investigate whether ChatGPT rated the scenarios to the same extent as do humans. Despite showing a clear trend that both types of raters provided consistent ratings of the items in accordance with the direction of the biased interpretation, Lin et al.\u0026rsquo;s study was underpowered, making it difficult to interpret their results. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eThe Present Study \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe overarching goal of the present study is to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials. By doing so, the overall cost of creating CBM-I materials may be significantly reduced and efficiency increased, thereby promoting its implementation on a broader scale and making it more accessible to both clinicians and clients as a treatment option for depression. We examined whether Copilot (https://copilot.microsoft.com), Microsoft\u0026rsquo;s state-of-the-art generative AI chatbot powered by the latest version of ChatGPT (v4.0), is capable of creating scenarios that capture situations and interpretation bias commonly associated with human experiences of depression. There is yet to be a study that has investigated the use of Copilot in creating CBM-I training materials; hence, this study is exploratory with no specific hypothesis. The study has the following objective: to compare participants\u0026rsquo; ratings of human-generated and AI-generated \u003cem\u003eraw\u003c/em\u003e and \u003cem\u003eCBM-I\u003c/em\u003e materials, which included scenarios with a negative and positive interpretation.\u003c/p\u003e"},{"header":"Method","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n \u003ch2\u003eProcedure\u003c/h2\u003e\n \u003cp\u003eThis study received institutional ethical approval (22/140). The present study included two stages which follows common development process of CBM-I training materials outlined in Hsu et al. (\u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e). Stage I involved the development of raw materials; Stage II involved adapting raw materials into CBM-I format (see Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e for a schematic outline of this process). This process involved using both human participants and Copilot to create and evaluate raw and CBM-I materials. Materials were validated on a set of predefined criteria:\u003c/p\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026bull; relevance of scenarios in reflecting experiences common to individuals with depression\u003c/p\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n \u003cp\u003e\u0026bull; severity of the interpretation of scenarios\u003c/p\u003e\n \u003cp\u003e\u0026bull;\u0026nbsp;\u003cem\u003ereadability of scenarios\u003c/em\u003e (for CBM-I materials only)\u003c/p\u003e\n \u003cp\u003eThese criteria were selected to assess whether scenarios are content specific to individuals with depression (\u003cem\u003erelevance\u003c/em\u003e) and that the interpretations of each scenario effectively capture a core cognitive feature of depression\u0026mdash;distorted negative thinking patterns (\u003cem\u003eseverity\u003c/em\u003e). Training materials were assessed for the level of clarity and comprehensiveness of scenarios (\u003cem\u003ereadability\u003c/em\u003e).\u003c/p\u003e\n \u003cp\u003eEach rating score is benchmarked against ratings of human-generated scenarios, which people with living/lived experience of depression creates. A deviation from human ratings in any direction on the rating scale would attend to our research question\u0026mdash;that AI-generated materials are not equivalent to human-generated materials. By following the materials development process of CBM-I outlined in Hsu et al., we can effectively evaluate the potential of generative AI in replicating the development of CBM-I content.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n \u003ch2\u003eParticipants\u003c/h2\u003e\n \u003cp\u003eParticipants in both stages of the study were adults with self-report of living/lived experience of depression (advertised as \u0026ldquo;adults with first-hand experience of depression\u0026rdquo;). As this is a non-clinical study, we did not specifically recruit participants with a formal diagnosis of depression nor did we include screening measures of depression; instead, participants were recruited from the community in New Zealand through social media pages, a local job seeking website, and flyer advertisements posted in various public domains, including hospitals, university campus, and community clinics. The number of participants that entered the study varied across the different stages of the study. Prior to rating, all participants completed the Hospital Anxiety and Depression Scale (HADS; Zigmond \u0026amp; Snaith, \u003cspan class=\"CitationRef\"\u003e1983\u003c/span\u003e) to measure their current level of depression and anxiety for the purpose of providing a description of participant mood at the time of the study. The HADS is a 14-item self-report questionnaire designed to measure the severity of depression and anxiety using a 4-point Likert scale, with final scores ranging from 0\u0026ndash;21 for each scale and grouped according to the degree of severity (normal\u0026thinsp;=\u0026thinsp;0\u0026ndash;7; borderline\u0026thinsp;=\u0026thinsp;8\u0026ndash;10; abnormal\u0026thinsp;=\u0026thinsp;11+). See Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e for a description of the raters.\u003c/p\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003e\u003cem\u003eRaters\u0026rsquo; Characteristics\u003c/em\u003e\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRater: Raw Materials (n\u0026thinsp;=\u0026thinsp;30)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRater: CBM-I Materials (n\u0026thinsp;=\u0026thinsp;30)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDepression Duration [Mean Years (SD)]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.80 (\u003cem\u003e7.90\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e11 (\u003cem\u003e8.78\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHADS Score\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDepression [Mean score (\u003cem\u003eSD)\u003c/em\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12.70 (\u003cem\u003e3.26\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e10.13 (\u003cem\u003e5.16\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eAnxiety [Mean score (\u003cem\u003eSD)\u003c/em\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.03 (\u003cem\u003e3.74\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e13.43 (\u003cem\u003e4.16\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eGender (female:male:other)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e23:3:4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e26:3:1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eEthnicity (NZ European:Māori:other)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e18:5:11\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e20:4:13\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eAge (Years) Mean (\u003cem\u003eSD\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e26.47 (\u003cem\u003e9.00\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e30.38 (\u003cem\u003e8.66\u003c/em\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\"\u003e\u003cem\u003eNote.\u003c/em\u003e The ethnicity ratio does not equate to N\u0026thinsp;=\u0026thinsp;30 because some raters identified with one or more ethnicity group.\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003c/p\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eStage I: Raw Materials\u003c/h3\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n \u003ch2\u003eRaw Materials Creation\u003c/h2\u003e\n \u003cp\u003eAs a part of another study on CBM-I targeting depression, experts with living/lived experience of depression (N\u0026thinsp;=\u0026thinsp;13) independently created a set of 130 raw items to describe their common everyday experiences (see Supplementary material, Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). Instructions given to the participants were \u003cstrong\u003e\u0026ldquo;\u003c/strong\u003e\u003cem\u003ewrite down 10 situations that could provoke depressive/negative interpretation; also provide a neutral/positive interpretation of that situation. The situation could be based on personal experience, someone else\u0026rsquo;s experience, or a situation that you can imagine experiencing.\u0026rdquo;\u003c/em\u003e Here is an example scenario: \u003cem\u003eA friend hasn\u0026rsquo;t replied to a text\u003c/em\u003e, and its interpretations: \u003cem\u003eThey\u0026rsquo;re ignoring me because they don\u0026rsquo;t like me\u003c/em\u003e (negative); \u003cem\u003ethey might be busy and will reply when they have time\u003c/em\u003e (neutral/positive). The negative interpretation is created only to illustrate the ambiguity of the scenario and support its face validity\u0026mdash;it is not used in the CBM-I training itself. The research team quasi-randomly selected 50 out of 130 raw items to be included in the present study, with 3 to 4 scenarios coming from each participant; we adopted a simple random sampling method using a random number generator to select scenarios that were created by each participant. This approach of selecting items maximized the variety of items to reflect depressive experiences that form part of the CBM-I intervention.\u003c/p\u003e\n \u003cp\u003eCopilot was prompted to generate another set of 50 items to reflect human experiences of depression (see Supplementary Material, Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Copilot has three conversation style options: \u0026lsquo;Precise, \u0026lsquo;Creative\u0026rsquo;, and \u0026lsquo;Balanced\u0026rsquo;. The \u0026lsquo;Precise\u0026rsquo; conversation style was selected to optimize the precision of scenarios to human experience of depression. The instructions given to Copilot were similar to the instructions provided to human participants: \u003cstrong\u003e\u0026ldquo;\u003c/strong\u003e\u003cem\u003eCreate different situations that someone with depression would commonly encounter and provide a depressive/negative interpretation and a neutral/positive interpretation of the situation\u0026rdquo;.\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eTen emotionally neutral items were created by the research team and included in the rating phase to serve as a manipulation check to ensure that raters are responding consistently and as expected. Manipulation check items were trivial statements designed to be unrelated to any emotional content (e.g., \u003cem\u003ePhones, first appearing in 1849, have evolved over time\u003c/em\u003e) presented with two sentences to complete the statement (e.g., \u003cem\u003eYou\u0026rsquo;re discussing the history of communication devices; You\u0026rsquo;re exploring the evolution of technology\u003c/em\u003e). Since these items are emotionally neutral, the expectation is that raters should give similar ratings to both sentences. Hence, any differences in ratings of check items may suggest the presence of confounding variables in influencing the overall ratings.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003eRaw Item Rating\u003c/h2\u003e\n \u003cp\u003eWhen developing these materials, it is common practice to refine content by working with living/lived experience experts to rate training content based on a set of pre-defined criteria. During the rating phase, all raw items were interleaved with 10 manipulation check items, yielding a total of 110 items. Scenarios were presented first, followed by two interpretations, one negative and one positive, that provided a different explanation to each scenario. Here is an example set that includes the situation along with the two interpretations:\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Taba\" border=\"1\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eSituation\u003c/strong\u003e: A friend hasn\u0026rsquo;t replied to a text.\u003c/p\u003e\n \u003cp\u003eHere are the two possible interpretations of this situation:\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation 1\u003c/strong\u003e: They\u0026rsquo;re ignoring me because they don\u0026rsquo;t like me.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation 2\u003c/strong\u003e: They might be busy and will reply when they have time.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eThe interpretations were presented in random order, and the order was blinded to the raters. The raters (N\u0026thinsp;=\u0026thinsp;30) rated all items independently and remotely on Qualtrics\u0026mdash;an online survey platform. A priori power calculation using G*Power3.1 (Faul et al., 2009) revealed\u0026thinsp;\u0026gt;\u0026thinsp;80% power at \u0026alpha;\u0026thinsp;=\u0026thinsp;.05 using an equivalence bound of \u003cem\u003ed\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.718 (Lakens, 2017). The equivalence bound was derived from earlier work that compared ChatGPT ratings to human ratings of scenarios used in the Similarity Rating Task (see Lin et al., \u003cspan class=\"CitationRef\"\u003e2023\u003c/span\u003e). Items were rated against the pre-defined rating criteria using a 7-point scale: \u0026lsquo;relevance\u0026rsquo; and \u0026lsquo;severity\u0026rsquo;. A higher rating on the \u0026lsquo;relevance\u0026rsquo; scale is defined as the item being more representative of the experiences of individuals with depression (i.e., 1\u0026thinsp;=\u0026thinsp;the least relevant to 7\u0026thinsp;=\u0026thinsp;the most relevant). A higher rating on the \u0026lsquo;severity\u0026rsquo; scale is defined as the item being more representative of the experiences of individuals with more severe depression (i.e., 1\u0026thinsp;=\u0026thinsp;the least negative to 7\u0026thinsp;=\u0026thinsp;the most negative).\u003c/p\u003e\n \u003cp\u003eEach item includes the situation and one of the two interpretations.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003eStage II: CBM-I Items\u003c/h2\u003e\n \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\n \u003ch2\u003eCBM-I Item Adaptation\u003c/h2\u003e\n \u003cp\u003eThe human-generated raw items were adapted into CBM-I standard format by the research team; The AI-generated raw items were adapted into CBM-I format by Copilot (see Supplementary Material, Table\u0026nbsp;3; Table\u0026nbsp;4; Table\u0026nbsp;5). For the manipulation check items, half were adapted into CBM-I format by researchers, and the other half were adapted by Copilot. Here is an example of an adapted CBM-I scenario: \u003cem\u003eYou send a text message to a friend. They have not replied yet. You think they are\u0026hellip;.\u003c/em\u003eand its interpretations: \u003cem\u003eIgnoring you\u003c/em\u003e (negative); \u003cem\u003eBusy\u003c/em\u003e (neutral/positive). Again, the negative interpretation is created only to illustrate the ambiguity of the scenario and support its face validity. The \u0026lsquo;Precise\u0026rsquo; conversation style of Copilot was selected. Instructions provided to Copilot were \u0026ldquo;\u003cem\u003eCognitive Bias Modification (CBM) format is 3 sentences long with the final word resolving the scenario in a negative manner or a positive manner. In CBM format, the interpretations are preferably 1 word but can be 2 or 3 if needed to make the sentence more readable.\u0026rdquo;\u003c/em\u003e An example of CBM-I format was also provided: \u0026ldquo;\u003cem\u003eSituation: You make a mistake at work. You receive some feedback about it. You think you are\u0026hellip;Negative Interpretation: incompetent; Neutral/Positive Interpretation: improving\u0026rdquo;\u003c/em\u003e along and with additional instructions to support this process (e.g., \u0026ldquo;\u003cem\u003eCBM format keeps all three sentences identical except the last word is changed. make the CBM format like this:\u0026rdquo; and the situation should be in three sentences\u0026rdquo;\u003c/em\u003e).\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n \u003ch2\u003eCBM-I Item Rating\u003c/h2\u003e\n \u003cp\u003eSimilar to the rating process for raw items, all CBM-I items were interleaved with 10 manipulation check items, yielding a total of 110 items. Scenarios were presented first, followed by two interpretations, one negative and one positive, that completed each scenario.\u003c/p\u003e\n \u003cp\u003eFor example:\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tabb\" border=\"1\" class=\"fr-table-selection-hover\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eSituation\u003c/strong\u003e: You send a text message to a friend. They have not replied yet. You think they are\u0026hellip;\u003c/p\u003e\n \u003cp\u003eHere are the two possible interpretations of this situation:\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation 1\u003c/strong\u003e: Ignoring you\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eInterpretation 2\u003c/strong\u003e: Busy\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eThe interpretations were presented in random order, and the order was blinded to the raters. Another group of 30 raters were recruited from the same sources using the same method as the raters from the raw items rating phase (see Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e for raters\u0026rsquo; characteristics). Raters rated all items independently and remotely on Qualtrics. Items were rated against the pre-defined rating criteria using a 7-point scale: \u0026lsquo;relevance\u0026rsquo; and \u0026lsquo;severity\u0026rsquo;. An additional \u0026lsquo;\u003cem\u003ereadability\u003c/em\u003e\u0026rsquo; criterion was included. A higher rating on this criterion reflected more clear and comprehensible scenarios.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eGiven the exploratory nature of the present study, we conducted equivalence testing using the Two One-Sided Test (TOST) for paired data (Schuirmann, \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e1987\u003c/span\u003e) and null hypothesis significance testing (NHST) to analyze content ratings. Equivalence testing included an equivalence bound of Cohen\u0026rsquo;s \u003cem\u003ed\u003c/em\u003e = -0.718 \u003cem\u003eto d\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.718 with a 90% confidence interval (see Lin et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Research data are available here: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://tinyurl.com/GENAICBM\u003c/span\u003e\u003cspan address=\"https://tinyurl.com/GENAICBM\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. For analyses, we reversed the scale so that 1\u0026thinsp;=\u0026thinsp;\u003cem\u003ethe least relevant to depression and most positive\u003c/em\u003e to 7\u0026thinsp;=\u0026thinsp;\u003cem\u003ethe most relevant to depression and negative\u003c/em\u003e. This is to allow for more intuitive interpretation of rating scores.\u003c/p\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003eStage I: Raw Items\u003c/h2\u003e\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e left panel, results of TOST for paired data showed that human-generated and Copilot-generated raw items were statistically non-equivalent across all rating criteria (see Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003ea for statistics). NHST using paired t-tests revealed convergent results in that ratings of human-generated and Copilot-generated raw items were significantly different across all rating criteria, with a difference in mean ratings falling between 0.1\u0026ndash;0.48 (see Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). Statistically equivalent ratings were found for manipulation check items, suggesting that confounding effects of the order of presentation of scenarios with a positive and negative interpretation were likely controlled.\u003c/p\u003e\u003cp\u003eDespite these findings, the direction of participants\u0026rsquo; ratings aligned with the emotional valence of the scenarios. Specifically, both AI- and human-generated negative scenarios were rated above the midpoint on the criteria \u0026lsquo;relevance\u0026rsquo; (M\u003csub\u003eAI-generated\u003c/sub\u003e = 5.93; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 5.63) and \u0026lsquo;severity\u0026rsquo; (M\u003csub\u003eAI-generated\u003c/sub\u003e = 6.06; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 5.89), leaning toward the higher end. This suggests consistency in how both sources generated excerpts that conveyed negative content. Similarly, positive scenarios were rated closer to the lower end of the scale (relevance: M\u003csub\u003eAI-generated\u003c/sub\u003e = 2.91; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 3.39 and severity: M\u003csub\u003eAI-generated\u003c/sub\u003e = 2.45; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 2.76), reflecting their intended positive valence. Furthermore, ratings of manipulation check items clustered around the midpoint (relevance: M\u003csub\u003eAI-generated\u003c/sub\u003e = 4.61; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 4.54 and severity: M\u003csub\u003eAI-generated\u003c/sub\u003e = 3.63; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 3.56),further supporting the overall consistency in rating directionality.\u003c/p\u003e\u003cp\u003eA careful examination of the ratings revealed that participants consistently rated the AI-generated items as more negative on the severity dimension and more relevant to common experiences of individuals with depression. A similar pattern of directionality was observed for positively valence items, with raters providing lower ratings on the AI-generated items. These results suggest that participants tended to rate AI-generated items were higher when they depicted negative content and lower when they depicted positive content, indicating a consistent directional trend in their evaluations.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003e\u003cb\u003ea\u003c/b\u003e \u003cem\u003eResults and Statistical Tests\u003c/em\u003e of Raw Items [Mean (SD), 90% confidence interval CI]\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRating Criteria\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAI-Generated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman-Generated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEquivalence Test\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDifference Test\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRelevance\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNegative Scenarios\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5.93\u003c/p\u003e\u003cp\u003e(0.77)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5.63\u003c/p\u003e\u003cp\u003e(0.70)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;0.98, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.83\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.41 to -0.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -4.91, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI 0.20 to 0.41\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePositive Scenarios\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e2.91\u003c/p\u003e\u003cp\u003e(1.02)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.39\u003c/p\u003e\u003cp\u003e(0.97)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -3.15, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.10\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI 0.37 to 0.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;7.08, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI -0.6 to -0.37\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eManipulation check items\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e4.61\u003c/p\u003e\u003cp\u003e(0.87)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4.54\u003c/p\u003e\u003cp\u003e(0.84)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -3.10, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.002\u003c/em\u003e, CI -0.2 to 0.07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -0.83, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.41,\u003c/p\u003e\u003cp\u003eCI-0.07 to -0.20\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSeverity\u003c/b\u003e:\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNegative interpretation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e6.06\u003c/p\u003e\u003cp\u003e(0.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5.89\u003c/p\u003e\u003cp\u003e(0.73)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -0.37, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.36\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.24 to -0.09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -3.56, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI 0.09 to -0.24\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePositive interpretation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e2.45\u003c/p\u003e\u003cp\u003e(0.90)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.76\u003c/p\u003e\u003cp\u003e(0.70)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -0.92, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.82\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI 0.2 to 0.41\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;4.85, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI -0.41 to -0.2\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eManipulation check items\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e3.63\u003c/p\u003e\u003cp\u003e(0.52)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3.56\u003c/p\u003e\u003cp\u003e(0.53)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -2.71, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.006\u003c/em\u003e, CI -0.18 to 0.03\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;0.23, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.23,\u003c/p\u003e\u003cp\u003eCI -0.03 to 0.18\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003eStage II: CBM-I Items\u003c/h2\u003e\u003cp\u003eFollowing the adaptation of raw items into CBM-I format, we examined both the equivalence and differences in mean ratings of human-adapted and Copilot-adapted CBM-I items. The same statistical analyses used in comparing raw items were performed for this comparison (see Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, right panel). With the exception of severity ratings, comparing human-adapted and Copilot-adapted CBM-I items resulted in statistically non-equivalent (and statistically different) ratings, with a difference in mean ratings falling between 0.03\u0026ndash;0.41 (see Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e\u003cp\u003eSimilar to the ratings of the raw items, participants\u0026rsquo; ratings for the relevance criterion\u0026mdash;but not severity\u0026mdash;generally aligned with the emotional content of both the negative and positive CBM-I scenarios and the manipulation check items, indicating consistent directionality in how emotional content was perceived. Participants also consistently rated the AI-generated CBM-I items as more relevant to common experiences of individuals with depression (M\u003csub\u003eAI-generated\u003c/sub\u003e = 5.93; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 5.51). A similar pattern of directionality was observed for positively valence items, with raters providing lower ratings on the AI-generated CBM-I items (M\u003csub\u003eAI-generated\u003c/sub\u003e = 3.45; M\u003csub\u003ehuman-generated\u003c/sub\u003e = 3.68).\u003c/p\u003e\u003cp\u003eFinally, statistically equivalent ratings were observed for manipulation check items only on the relevance dimension only\u0026mdash;not severity. This result suggests that the order in which positive and negative interpretations were presented may have influenced participants\u0026rsquo; severity ratings, potentially explaining the inconsistencies in ratings we found between the relevance and severity scales.\u003c/p\u003e\u003cp\u003eTaken together, mean ratings of both Copilot-generated and human-generated raw items and adapted CBM-I items were both statistically non-equivalent and different, with the exception of the severity ratings on CBM-I items. The overall direction of ratings between AI-and human-generated scenarios were the same and consistent with the scenarios\u0026rsquo; emotional content.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003e\u003cb\u003eb\u003c/b\u003e \u003cem\u003eResults and Statistical Tests\u003c/em\u003e of CBM-I Items [Mean (SD), 90% confidence interval CI]\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRating Criteria\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAI-Generated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHuman-Generated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEquivalence Test\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDifference Test\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRelevance:\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNegative Scenarios\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5.93 (0.77)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e5.51 (0.84)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;2.79, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.10\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.52 to -0.31\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -6.72, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI 0.31 to 0.52\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePositive Scenarios\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e3.45 (1.08)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e3.68 (1.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -0.29, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.61\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI 0.13 to 0.31\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;4.22, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001,\u003c/p\u003e\u003cp\u003eCI -0.31 to -0.13\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eManipulation check items\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e3.94 (0.68)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e4.04 (0.60)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;2.96, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.003\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.08 to 0.28\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;0.97, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.34,\u003c/p\u003e\u003cp\u003eCI -0.28 to 0.08\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSeverity\u003c/b\u003e:\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNegative interpretation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5.62 (0.93)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e5.65 (0.95)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;3.18, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.002\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.05 to 0.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;0.75, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.46,\u003c/p\u003e\u003cp\u003eCI -0.12 to 0.045\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePositive interpretation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e3.10 (0.77)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e3.02 (0.83)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29) = -2.42, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.01\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.17 to 0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -1.51, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.14,\u003c/p\u003e\u003cp\u003eCI -0.01 to 0.17\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eManipulation check items\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e4.18 (0.62)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e3.92 (0.44)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;.81, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.21\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI \u0026minus;\u0026thinsp;0.41 to -0.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29) = -3.13, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.004,\u003c/p\u003e\u003cp\u003eCI 0.12 to 0.41\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eReadability\u003c/b\u003e:\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5.91 (1.47)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e5.87 (1.47)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;1.49, \u003cem\u003ep\u0026thinsp;=\u0026thinsp;.07\u003c/em\u003e,\u003c/p\u003e\u003cp\u003eCI -0.08 to -0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003et(29)\u0026thinsp;=\u0026thinsp;2.44, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;.02,\u003c/p\u003e\u003cp\u003eCI 0.01 to 0.08\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eCognitive bias modification-interpretation (CBM-I) is a digital therapeutic that utilizes ambiguous text-based scenarios to capture and modify damaging interpretation bias known to contribute to depression. An integral part of CBM-I\u0026rsquo;s success in modifying bias is the quality of training scenarios. The development process of these key materials is often onerous and costly, involving an iterative process of extensive input from multiple stakeholders to create and refine training content. The aim of the study was to examine whether materials created by generative artificial intelligence (AI) are equivalent to human-generated materials, as per pre-defined rating criteria. We followed a typical development procedure, as outlined in Hsu et al. (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), to create CBM-I items; first, developing raw items and then adapting the raw items into CBM-I format. Items were generated by AI (CoPilot) and human experts (i.e., people with living/lived experience of depression). We compared experts\u0026rsquo; ratings of AI-generated and human-generated items.\u003c/p\u003e\u003cp\u003eWith a larger sample size and using a more powerful AI platform compared to ChatGPT3.5 (Wang et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), our results offer a new perspective on Lin et al.\u0026rsquo;s (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) underpowered study where the authors used GPT3.5 to rate scenarios presented in the Similarity Rating Task, a bias assessment that is similar to the format of CBM-I. In our study, we found statistically non-equivalent and significant differences in mean ratings between human- and AI-generated raw and CBM-I items. There is a paucity of studies on generative AI in creating mental health scenarios, and the closest comparison we could draw from are studies on ChatGPT\u0026rsquo;s capabilities in generating clinical vignettes for medical education. In this area of research, concerns have been raised regarding ChatGPT\u0026rsquo;s accuracy and reliability in producing clinical scenarios that accurately depict common medical presentations (Alam et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), possibly due to limitations in its training data (Sahu et al., \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). In a scoping review, Xu et al. (\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) found that, in 42.5% of the reviewed articles, researchers have discussed that ChatGPT may present incorrect or inconsistent medical information and concepts in AI-generated clinical vignettes.\u003c/p\u003e\u003cp\u003eDespite our data indicating statistical non-equivalence in the overall ratings of human-generated and AI-generated materials, several important caveats must be considered when interpreting these findings. First, our study found no statistical differences in participants\u0026rsquo; ratings of severity of CBM-I items; instead, statistical equivalence was established. One possible explanation for this result is that the manipulation check items failed to show equivalence, suggesting that confounding variables may have influenced how participants rated the items. Next, caution is needed when generalizing results from studies that used AI-generated clinical vignettes, as the vignettes may be fundamentally different from people\u0026rsquo;s experience of depression and serve different purposes. Furthermore, many of these studies adopted a weaker study design, such as comparing participants\u0026rsquo; ratings of excerpts rather than using other more rigorous methods like randomized control trials (cf. Coşkun et al., \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Similar to our study, ratings are often based on a limited set of criteria used to evaluate the similarity between AI-generated and human-generated materials. However, other factors\u0026mdash;such as accuracy, readability, and clinical outcomes\u0026mdash;as well as broader considerations, including safety and societal implications, are also crucial in assessing the effectiveness, reliability, and quality of AI-generated materials. These dimensions should be included in future avenues of study.\u003c/p\u003e\u003cp\u003eFinally, when attending to our research question on whether AI-generated CBM-I training materials are equivalent to human-generated materials, it is important to consider that deviations from ratings of human-generated materials\u0026mdash;the benchmark in our study\u0026mdash;do not necessarily undermine the usefulness of AI-generated content (Lin et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). This is particularly relevant when AI- and human-generated excerpts show similar rating patterns that align with the intended emotional content. In our study, both sources generated excerpts that were either consistent with negative or positive scenarios. In most cases, participants also consistently rated AI-generated items higher when they depicted negative content and lower when they depicted positive content, indicating a consistent directional trend in their evaluations.\u003c/p\u003e\u003cp\u003eMoreover, it is important to examine results beyond statistical testing and consider the definition of equivalence in a clinical setting. For instance, by considering actual differences in mean ratings of materials, our data showed that the mean ratings fell within one unit/point on a 7-point rating scale between 0.03\u0026ndash;0.41, suggesting that the differences were minimal and may have limited impact clinically. More specifically, our data on the relevance criterion showed that mean ratings for both AI-generated items (M\u0026thinsp;=\u0026thinsp;5.93) and human-generated items (M\u0026thinsp;=\u0026thinsp;5.51) were above the mid-point for negative interpretations; ratings of positive interpretation were below the mid-point (M\u003csub\u003eAI\u0026minus;generated\u003c/sub\u003e = 3.45; M\u003csub\u003ehuman\u0026minus;generated\u003c/sub\u003e = 3.68). In a medical education study, Benoit (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) demonstrated that, despite ChatGPT creating clinical vignettes that included additional disease symptoms that was beyond what it was instructed to provide, the extraneous information depicted an accurate clinical symptom.\u003c/p\u003e\u003cp\u003eTaken together, despite the statistically non-equivalence of participants\u0026rsquo; ratings of human-generated and AI-generated materials, when interpreting the data, researchers should consider other confounding variables and evaluate what \u0026lsquo;true equivalence\u0026rsquo; really means. Future directions of research should consider using more rigorous study designs\u0026mdash;a randomized control trial\u0026mdash;and aim to examine the implementation of AI-generated items for CBM-I against human-generated items. Furthermore, carefully analyzing the excerpts and comparing them qualitatively may provide more rich understanding of AI-generated items. Finally, given the rapid advancement of generative AI technology, more recent or future versions of AI chatbots may have enhanced capabilities of generating scenarios that closely align with human experiences of depression.\u003c/p\u003e\u003cp\u003eThe benefits of using generative AI chatbots to create CBM-I training materials are beyond that of automation. Generative AI chatbots uses specialized algorithms to seek out patterns from human prompts and uses reinforcement learning and reward models to provide a response. The implication of this would mean individualised CBM-I training materials tailored to individual service users (Ahmad et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), which may improve the content specificity of materials and promote better bias modification effects (Mackintosh et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Savulich et al., \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Vancleef \u0026amp; Peters, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2008\u003c/span\u003e). What this may look like practically warrants further investigation and encouraged in future avenue of studies. It could mean that the chatbot is trained using individualized data to create similar descriptions of the individual\u0026rsquo;s experiences and biases.\u003c/p\u003e\u003cp\u003eThere are some limitations to the present study that should be considered in future studies. One major issue is the representation of the scenarios. That is, the validity of generalizing experiences from 13 individuals with depression may be low. In future studies, it is encouraged that researchers gather everyday experiences from a larger and diverse population, considering age, gender, ethnicity, and years of depression. An additional limitation is that only simple prompting was used to instruct Copilot in creating scenarios. Prompts in AI are pivotal for optimizing outputs (Bozkurt, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) and several strategies have been researched and developed (Liu et al., \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). It is important to test different prompt strategies to determine the optimal prompt for creating CBM-I training items. Finally, although the use of AI-generated scenarios may reduce cost associated with developing CBM-I materials, one should consider other potential cost of using AI, such as energy and environmental cost. A cost-benefit analysis may be warranted in future studies.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn conclusion, this study provides initial insights into the comparison between human-generated and AI-generated CBM-I materials. While the findings revealed AI- and human-generated scenarios are statistically non-equivalent, as per pre-defined rating criteria, it is important to recognize the potential of AI-generated materials on a more practical level. On this basis, our findings should be viewed as part of a broader and ongoing research in AI\u0026rsquo;s capabilities of generating excerpts for clinical use. Future research with a larger sample size and clinical validation would be necessary to better understand the effectiveness of AI in this context. Using generative AI in developing treatment materials for CBM-I training for depression could significantly reduce overall costs and increase the efficiency of the materials development process. This, in turn, could facilitate the broader dissemination and implementation of an effective intervention for depression.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAhmad, R., Siemon, D., Gnewuch, U., \u0026amp; Robra-Bissantz, S. (2022). Designing personality-adaptive conversational agents for mental health care. \u003cem\u003eInformation Systems Frontiers, \u003c/em\u003e\u003cem\u003e24\u003c/em\u003e(3), 923\u0026ndash;43. https://doi.org/10.1007/s10796-022-10254-9\u003c/li\u003e\n\u003cli\u003eAlam, F., Lim, M. A., \u0026amp; Zulkipli, I. N. (2023). Integrating AI in medical education: embracing ethical usage and critical understanding. \u003cem\u003eFrontiers in Medicine, 10, article \u003c/em\u003e1279707. https://doi.org/10.3389/fmed.2023.1279707\u003c/li\u003e\n\u003cli\u003eAlanezi, F. (2024). Assessing the Effectiveness of ChatGPT in Delivering Mental Health Support: A Qualitative Study. \u003cem\u003eJournal of Multidisciplinary Healthcare, 17, \u003c/em\u003e461\u0026ndash;471. https://doi.org/10.2147/JMDH.S447368\u003c/li\u003e\n\u003cli\u003eBlackwell, S. E., Browning, M., Mathews, A., Pictet, A., Welch, J., Davies, J., et al. (2015). Positive imagery-based cognitive bias modification as a web-based treatment tool for depressed adults: A randomized controlled trial. \u003cem\u003eClinical Psychological Science, 3\u003c/em\u003e(1), 91\u0026ndash;111. https://doi.org/10.1177/2167702614560746\u003c/li\u003e\n\u003cli\u003eBeck, A. T. (1976). \u003cem\u003eCognitive therapy and the emotional disorders.\u003c/em\u003e International Universities Press.\u003c/li\u003e\n\u003cli\u003eBenoit, J. R. A. (2023). ChatGPT for clinical vignette generation, revision, and evaluation. \u003cem\u003emedRxiv. \u003c/em\u003ehttps://doi.org/10.1101/2023.02.04.23285478\u003c/li\u003e\n\u003cli\u003eBowler, J. O., Mackintosh, B., Dunn, B. D., Mathews, A., Dalgleish, T., \u0026amp; Hoppitt, L. (2012). A comparison of cognitive bias modification for interpretation and computerized cognitive behavior therapy. \u003cem\u003eJournal of Consulting and Clinical Psychology, 80\u003c/em\u003e(6), 1021\u0026ndash;1033. https://doi.org/0022-006X/12/$12.00\u003c/li\u003e\n\u003cli\u003eBozkurt (2024). Tell me your prompts and I will make them true: the alchemy of prompt engineering and generative AI. \u003cem\u003eOpen Praxis, 16\u003c/em\u003e(2), 111\u0026ndash;118. https://doi.org/10.55982/openpraxis.16.2.661\u003c/li\u003e\n\u003cli\u003eCarlbring, P., Apelstrand, M., Sehlin, H., Amir, N., Rousseau, A., Hofmann, S. G., et al. (2012). Internet-delivered attention bias modification training in individuals with social anxiety disorder: A double-blind randomized controlled trial. \u003cem\u003eBMC Psychiatry, 12\u003c/em\u003e, 66. https://doi.org/10.1186/1471-244X-12-66\u003c/li\u003e\n\u003cli\u003eChen J., Short, M., \u0026amp; Kemps, E. (2020). Interpretation bias in social anxiety: A systematic review and meta-analysis. \u003cem\u003eJournal of Affect Disorder, 276\u003c/em\u003e(3)\u003cem\u003e, \u003c/em\u003e1119\u0026ndash;1130. https://10.1016/j.jad.2020.07.121\u003c/li\u003e\n\u003cli\u003eCoskun, O., Kiyak, Y. S., \u0026amp; Budakoglu, I. I. (2024). ChatGPT to generate clinical vignettes for teaching and multiple-choice questions for assessment: A randomized controlled experiment. \u003cem\u003eMedical Teacher, 13, \u003c/em\u003e1\u0026ndash;7. https://doi.org/10.1080/0142159X.2024.2327477\u003c/li\u003e\n\u003cli\u003eDuong, D. \u0026amp; Solomon, B. D. (2023). Analysis of large-language model \u003cem\u003evs\u003c/em\u003e human performance for genetics questions. \u003cem\u003eEuropean Journal of Human Genetics, 32\u003c/em\u003e(4), 466\u0026ndash;468. https://doi.org/10.1038/s41431-023-01396-8\u003c/li\u003e\n\u003cli\u003eEveraert, J., Podina, I. R., \u0026amp; Koster, E. H. W. (2017). A comprehensive meta-analysis of interpretation biases in depression. \u003cem\u003eClinical Psychology Review, 58, \u003c/em\u003e33\u0026ndash;48. https://doi.org/10.1016/j.cpr.2017.09.005\u003c/li\u003e\n\u003cli\u003eGarety, P. A., Kuipers, E., Fowler, D., Freeman, D., \u0026amp; Bebbington, P. E. (2001). A cognitive model of the positive symptoms of psychosis. \u003cem\u003ePsychological Medicine\u003c/em\u003e, \u003cem\u003e31\u003c/em\u003e(2), 189\u0026ndash;195. https://doi.org/10.1017/s0033291701003312\u003c/li\u003e\n\u003cli\u003eHa, S.W. \u0026amp; Kim, J. (2020). Designing a scalable, accessible, and effective mobile app based solution for common mental health problems. \u003cem\u003eInternational Journal of Human-Computer Interaction,\u003c/em\u003e \u003cem\u003e36\u003c/em\u003e(35), 1354\u0026ndash;1367. https://doi.org/10.1080/10447318.2020.1750792\u003c/li\u003e\n\u003cli\u003eHallion, L. S., \u0026amp; Ruscio, A. M. (2011). A meta-analysis of the effect of cognitive bias modification on anxiety and depression. \u003cem\u003ePsychological Bulletin, 137\u003c/em\u003e(6), 940\u0026ndash;958. https://doi.org/10.1037/a0024355\u003c/li\u003e\n\u003cli\u003eHirsch, C. R., Krah\u0026eacute;, C., Whyte, J., Loizou, S., Bridge, L., Norton, S., \u0026amp; Mathews, A. (2018). Interpretation training to target repetitive negative thinking in generalized anxiety disorder and depression. \u003cem\u003eJournal of Consulting and Clinical Psychology, 86\u003c/em\u003e(12), 1017\u0026ndash;1030. https://doi.org/10.1037/ccp0000310\u003c/li\u003e\n\u003cli\u003eHsu, C. W. \u0026amp; Akuhatahuntington, Z. (2024). Implicit bias training for New Zealand medical students using cognitive bias modification: An outline of material development. \u003cem\u003eNew Zealand Journal of Psychology\u003c/em\u003e, \u003cem\u003e52\u003c/em\u003e(3), 30\u0026ndash;43.\u003c/li\u003e\n\u003cli\u003eHsu, C. W., Stahl, D., Mouchlianitis, E., Peters, E., Vamvakas, G., Keppens, J., Watson, M., Schmidt, N., Jacobsen, P., McGuire, P., Shergill, S., Kabir, T., Hirani, T., Yang, Z. Y., \u0026amp; Yiend, J. (2023). User-Centered Development of STOP (Successful Treatment for Paranoia): Material Development and Usability Testing for a Digital Therapeutic for Paranoia. \u003cem\u003eJMIR Human Factors, 10, \u003c/em\u003ee45453. https://doi.org/10.2196/45453\u003c/li\u003e\n\u003cli\u003eHuh, S. (2023). Are ChatGPT\u0026rsquo;s knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination? A descriptive study. \u003cem\u003eJournal of Educational Evaluation of Health Professions, 20\u003c/em\u003e, 1. https://doi.org/10.3352/jee\u0026shy;hp.2023.20.1\u003c/li\u003e\n\u003cli\u003eJeyaraman, M., Ramasubramanian, S., Balaji, S., Jeyaraman, N., Nallakumarasamy, A., \u0026amp; Sharma, S. (2023). ChatGPT in action: Harnessing artificial intelligence potential and addressing ethical challenges in medicine, education, and scientific research. \u003cem\u003eWorld Journal of Methodology, 13\u003c/em\u003e(4), 170\u0026ndash;178. https://doi.org/10.5662/wjm.v13.i4.170 \u003c/li\u003e\n\u003cli\u003eLakens, D. (2018). Equivalence Testing for Psychological Research: A Tutorial. \u003cem\u003eAdvances in Methods and Practices in Psychological Science, 1\u003c/em\u003e(2), 259\u0026ndash;269. https://doi.org/10.1177/2515245918770963\u003c/li\u003e\n\u003cli\u003eLi J. W., Ma, H., Yang, H., Yu, H., \u0026amp; Zhang, N. (2022). Cognitive bias modification for adult\u0026rsquo;s depression: A systematic review and meta-analysis. \u003cem\u003eFrontiers in Psychology, 13, \u003c/em\u003e968638. https://10.3389/fpsyg.2022.968638\u003c/li\u003e\n\u003cli\u003eLin, C., C., Akuhata-Huntington, Z., \u0026amp; Hsu, C. W. (2023). Comparing ChatGPT\u0026rsquo;s ability to rate the degree of stereotypes and the consistency of stereotype attribution with those of medical students in New Zealand in developing a similarity rating test: a methodological study. \u003cem\u003eJournal of Educational Evaluation of Health Professions, 20\u003c/em\u003e, 17. https://10.3352/jeehp.2023.20.17\u003c/li\u003e\n\u003cli\u003eLin, C. C., du Plooy, K., Gray, A., Brown, D., Hobbs, L., Patterson, T., Tan, V., Fridberg, D., Hsu, C. W. (2024). The performance of ChatGPT on short-answer questions in a psychiatry examination: A pilot study. \u003cem\u003eTaiwanese Journal of Psychiatry, 38\u003c/em\u003e(2), 94\u0026ndash;98. https://doi.org/10.4109/TPSY.TPSY_19_24\u003c/li\u003e\n\u003cli\u003eLiu, H., Li, X., Han, B, \u0026amp; Liu, X. (2017). Effects of cognitive bias modification on social anxiety: A meta-analysis. \u003cem\u003ePLoS One, 12\u003c/em\u003e(4): e0175107\u003c/li\u003e\n\u003cli\u003eLiu, P. F., Y, W. Z., Fu, J. L., Jiang, Z. B., Hayashi, H., \u0026amp; Neubig, G. (2023). Pre-train, prompt, and predict. A systematic survey of prompting methods in natural language processing. \u003cem\u003eACM Computing Surveys, 55\u003c/em\u003e(9), 1\u0026ndash;35. https://doi.org/10.1145/3560815\u003c/li\u003e\n\u003cli\u003eMackintosh, B., Mathews, A., Eckstein, D., \u0026amp; Hoppitt, L. (2013). Specificity effects in the modification of interpretation bias and stress reactivity. \u003cem\u003eJournal of Experimental Psychopathology, 4\u003c/em\u003e(2)\u003cem\u003e, \u003c/em\u003e133\u0026ndash;147\u003cu\u003e. \u003c/u\u003ehttps://doi.org/10.5127/jep.025711\u003c/li\u003e\n\u003cli\u003eMartinelli, A., Gr\u0026uuml;ll, J., \u0026amp; Baum, C. (2022). Attention and interpretation cognitive bias change: A systematic review and meta-analysis of bias modification paradigms. \u003cem\u003eBehaviour Research and Therapy, 157, \u003c/em\u003e104180. https://doi.org/10.1016/j.brat.2022\u003cem\u003e \u003c/em\u003e\u003c/li\u003e\n\u003cli\u003eMathews, A. \u0026amp; Mackintosh B. (2000). Induced emotional interpretation bias and anxiety. \u003cem\u003eJournal of Abnormal Psychology, \u003c/em\u003e\u003cem\u003e109\u003c/em\u003e(4), 602\u0026ndash;615. https://doi.org/10.1037/0021-843X.109.4.602\u003c/li\u003e\n\u003cli\u003eMontazeri, M., Galavi, Z., \u0026amp; Ahmadian, L. (2024). What are the applications of ChatGPT in healthcare: Gain or loss? \u003cem\u003eHealth Science Reports, 7\u003c/em\u003e(2), e1878. https://10.1002/hsr2.1878\u003c/li\u003e\n\u003cli\u003eRude, S., Valdez, C.R., Odom, S., \u0026amp; Ebrahimi, A. (2003). Negative cognitive biases predict subsequent depression. \u003cem\u003eCognitive Therapy and Research, \u003c/em\u003e\u003cem\u003e27\u003c/em\u003e(4), 415\u0026ndash;429. https://doi.org/10.1023/A:1025472413805\u003c/li\u003e\n\u003cli\u003eSahu, P. K., Benjamin, L. A., Singh, A. G., \u0026amp; Williams-Persad, A. (2023). ChatGPT in research and health professions education: challenges, opportunities, and future directions. \u003cem\u003ePostgraduate Medical Journal, 100\u003c/em\u003e(1179), 50\u0026ndash;55. https://doi.org/10.1093/postmj/qgad090\u003c/li\u003e\n\u003cli\u003eSavulich, G., Freeman, D., Shergill, S., \u0026amp; Yiend, J. (2015). Interpretation biases in paranoia. \u003cem\u003eBehavior Therapy, 46\u003c/em\u003e(1), 110\u0026ndash;124. https://doi.org/10.1016/j.beth.2014.08.002\u003c/li\u003e\n\u003cli\u003eSavulich, G., Shergill, S. S., \u0026amp; Yiend, J. (2017). Interpretation biases in clinical \u003c/li\u003e\n\u003cli\u003eparanoia. \u003cem\u003eClinical Psychological Science\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(6), 985\u0026ndash;1000. https://doi.org/10.1177/2167702617718180\u003c/li\u003e\n\u003cli\u003eSchuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. \u003cem\u003eJournal of Pharmacokinetics and Pharmacodynamics, 15\u003c/em\u003e(6), 657\u0026ndash;680. https://doi.org/10.1007/BF01068419\u003c/li\u003e\n\u003cli\u003eSteinman, S. A., Namaky, N., Toton, S. L., Meissel, E. E. E., John, A. T. S., Pham, N. H., Werntz, A., Valladares, T. L., Gorlin, E. I., Arbus, S., Beltzer, M., Soroka, A., \u0026amp; Teachman, B. A. (2021). Which variations of a brief cognitive bias modification session for interpretations lead to the strongest effects? \u003cem\u003eCognitive Therapy and Research, 45\u003c/em\u003e(2), 367\u0026ndash;382. https://doi.org/10.1007/s10608-020-10168-3\u003c/li\u003e\n\u003cli\u003eVancleef, L. M. G. \u0026amp; Peters, M. L. (2008). Examining content specificity of negative interpretation biases with the Body Sensations Interpretation Questionnaire (BSIQ). \u003cem\u003eJournal of Anxiety Disorder, 22\u003c/em\u003e(3), 401\u0026ndash;415. https://doi.org/10.1016/j.janxdis.2007.05.006. \u003c/li\u003e\n\u003cli\u003eWang, G., Gao, K., Liu, Q. Y., Wu, Y. X., Zhang, K. J., Zhou, W., \u0026amp; Gui, C. B. (2023). Potential and Limitations of ChatGPT 3.5 and 4.0 as a Source of COVID-19 Information: Comprehensive Comparative Analysis of Generative and Authoritative Information. \u003cem\u003eJournal of Medical Internet Research, 25, \u003c/em\u003ee49771. https://doi.org/10.2196/49771\u003c/li\u003e\n\u003cli\u003eXu, X. J., Chen, Y. X., \u0026amp; Miao, J. (2024). Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review. \u003cem\u003eJournal of Educational Evaluation for Health Professions, \u003c/em\u003e21(6). https://doi.org/10.3352/jeehp.2024.21.6\u003c/li\u003e\n\u003cli\u003eYiend, J., Lam, C. L. M., Schmidt, N., Crane, B., Heslin, M., Kabir, T., McGuire, P., Meek, C., Mouchlianitis, E., Peters, E., Stahl, D., Trotta, A., \u0026amp; Shergill, S. (2023). Cognitive bias modification for paranoia (CBM-pa): a randomised controlled feasibility study in patients with distressing paranoid beliefs. \u003cem\u003ePsychological Medicine, 53\u003c/em\u003e(10), 4614\u0026ndash;4626. https://doi.org/10.1017/S0033291722001520\u003c/li\u003e\n\u003cli\u003eYiend, J., Lee, J. S., Tekes, S., Atkins, L., Mathews, A., Vrinten, M., Ferragamo, C., \u0026amp; Shergill, S. (2013). Modifying Interpretation in a Clinically Depressed Sample Using \u0026lsquo;Cognitive Bias Modification-Errors\u0026rsquo;: A Double Blind Randomised Controlled Trial. \u003cem\u003eCognitive Therapy and Research, 38\u003c/em\u003e(2), 146\u0026ndash;159. https://10.1007/s10608-013-9571-y \u003c/li\u003e\n\u003cli\u003eZigmond, A. S., \u0026amp; Snaith, R. P. (1983). \u003cem\u003eThe hospital anxiety and depression scale. Acta Psychiatrica Scandinavica, 67(6), 361\u0026ndash;370. \u003c/em\u003ehttps://doi.org/10.1111/j.1600-0447.1983.tb09716.x\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Supplementary Material","content":"\u003cp\u003eSupplementary Tables 1-5 are not available with this version.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of Otago","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"cognitive bias modification, interpretation bias, depression, ChatGPT, generative AI chatbot, Copilot, digital mental health","lastPublishedDoi":"10.21203/rs.3.rs-7530420/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7530420/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePeople with depression tend to interpret ambiguous events in a negative biased direction, which contributes to symptomatology. Cognitive bias modification-interpretation (CBM-I) is a digital therapeutic that targets negative interpretation bias using text-based scenarios. CBM-I offers greater flexibility in combating depressive-related issues, but the development of its training materials can be costly. In the present study, we used Generative AI to produce CBM-I training materials and compared them with human-generated materials. The aim is to examine whether AI-generated materials are equivalent to human-generated materials, as per a set of pre-defined criteria designed to capture common experiences of individuals with depression. We followed the typical CBM-I materials development procedure, first creating raw items and then adapting them into standard CBM-I format. We compared participants\u0026rsquo; ratings of 100 raw scenarios and 100 CBM-I scenarios, half of which were created by Copilot and half created by people with depression. Living/lived experts of depression (N\u0026thinsp;=\u0026thinsp;30) rated raw items, and another 30 experts rated CBM-I items on readability, relevance, and severity of scenarios as related to depression. With the exception of severity ratings, results revealed that ratings of human-generated and AI-generated scenarios were statistically non-equivalent. The differences in the overall actual mean ratings, however, were small (range 0.03\u0026ndash;0.41); the overall direction of ratings between AI-and human-generated scenarios were the same and consistent with the scenarios\u0026rsquo; emotional content. Interpretation of the data and future implications of outsourcing AI in CBM-I materials production are discussed.\u003c/p\u003e","manuscriptTitle":"Generative artificial intelligence in mental health: A preliminary study on automating materials development for cognitive bias modification.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-09 11:24:14","doi":"10.21203/rs.3.rs-7530420/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7f2acb29-0cae-4bb8-b176-08c412927b07","owner":[],"postedDate":"September 9th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":54161855,"name":"Psychology"}],"tags":[],"updatedAt":"2025-09-09T11:24:14+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-09 11:24:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7530420","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7530420","identity":"rs-7530420","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00