Evaluating Evidence-Based Communication through Generative AI using a Cross-Sectional Study with Laypeople Seeking Screening Information | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Evaluating Evidence-Based Communication through Generative AI using a Cross-Sectional Study with Laypeople Seeking Screening Information Felix G. Rebitschek, Alessandra Carella, Silja Kohlrausch-Pazin, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6220209/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 09 Jun, 2025 Read the published version in npj Digital Medicine → Version 1 posted 10 You are reading this latest preprint version Abstract Large language models (LLMs) are used to seek health information. We investigate the prompt-dependent compliance of LLMs with evidence-based health communication guidelines and evaluate the efficacy of a minimal behavioral intervention for boosting laypeople’s prompting. Study 1 systematically varied prompt informedness, topic, and LLMs to evaluate LLM compliance. Study 2 randomized 300 UK participants to interact with LLMs under standard or boosted prompting conditions. Independent blinded raters assessed LLM response with 2 instruments. Study 1 found that LLMs failed evidence-based health communication standards, even with informed prompting. The quality of responses was found to be contingent upon prompt informedness. Study 2 revealed that laypeople frequently generated poor-quality responses; however, a simple boost improved response quality, though it remained below optimal standards. These findings underscore the inadequacy of LLMs as a standalone health communication tool. It is imperative to enhance LLM interfaces, integrate them with evidence-based frameworks, and teach prompt engineering. Study Registration : German Clinical Trials Register (DRKS) (Reg. No.: DRKS00035228) Ethical Approval : Ethics Committee of the University of Potsdam (Approval No. 52/2024) Scientific community and society/Scientific community/Education Health sciences/Diseases/Cancer/Breast cancer Health sciences/Diseases/Cancer/Cancer screening Health sciences/Diseases/Cancer/Urological cancer/Prostate cancer Health sciences/Health care/Health policy Health sciences/Health care/Health services Health sciences/Health care/Public health Scientific community and society/Social sciences/Decision making Large Language Models LLM Health Communication Evidence-Based Information quality Boosting Figures Figure 1 Figure 2 Introduction The internet has become a primary source of health information for people in Western countries. 1 However, many online resources fail to adhere to evidence-based communication standards. 2 As the practical implementation of guidelines for evidence-based health communication 3 , 4 remains challenging, much of the health information available online does not support informed health decision-making, it lacks necessary accuracy and rigor, and in some cases even spreads misinformation, particularly in areas such as cancer. 5 The rapid emergence of artificial intelligence (AI) tools, particularly large language models (LLMs) such as OpenAI's ChatGPT, Mistral AI's Le Chat, has opened novel possibilities for digital health communication – even in professional settings. 6 Laypeople are increasingly turning to these platforms to seek answers to health-related questions. 7 However, can LLM users reliably obtain correct information that complies with established health communication guidelines? To address this, we conduct two complementary studies aimed at evaluating and enhancing the quality of health communication provided by LLMs. Although LLMs are not designed for health communication, laypeople use them to gather health information. In a convenience online sample of adult panelists from February to March 2023, 7·2% reported regularly using LLMs for health-related topics—often in combination with search engines and online health communities. 8 This number increased significantly, with 73·6% of an online community convenience sample between December 2023 and January 2024 and 32·6% of a U.S. online panel in February 2024 using LLMs for health information. 9 Though often correct, LLM-generated responses to health-related queries show substantial variability. 10 , 11 The ‘accuracy’ or ‘reliability’ of LLM responses has been subject to numerous descriptive studies, including a recent systematic review of 88 patient education studies. 12 Yet, what does accuracy or reliability truly mean in this context? LLM-generated responses to health-related questions has been assessed in various ways, such as readability, 13 – 18 expert-rated appropriateness of factual explanations, 15 , 18 – 27 and agreement with health authorities and guidelines. 28 – 35 However, what remains missing, is the use of the best available evidence as the ground truth. LLM information and communication studies consistently lack a well-defined health-care premise (e.g. informed decision-making). This absence undermines both the criteria for response assessment (e.g. compliance with evidence-based communication standards) and the structured selection of prompts (e.g. aligning with information requirements). Moreover, these studies often fail to ground their assessments in the best available evidence, such as high-quality evidence synthesis. Evidence-based guidelines, such as those developed by the Working Group for Good Practice in Health Information (GPHI), provide a structured framework for health communication. 4 , 36 They emphasize, for instance, the importance of using numerical data to present risks and benefits based on the best available evidence. Regarding validated instruments for assessing information quality, at least three LLM studies 10 , 13 , 37 have applied the DISCERN standards, which from today's perspective, are considered outdated, while one study utilized a validated patient information quality instrument. 38 The primary aim of Study 1, therefore, is to assess the extent to which content generated by state-of-the-art LLMs—specifically OpenAI’s ChatGPT, Google Gemini, and Mistral AI—adheres to guideline-based criteria, with a particular focus on essential aspects of health risk communication for informed decision-making. These criteria encompass clear communication of benefits and harms, inclusion of reference class information, the appropriate presentation of numerical effects, the quality of evidence, and accurate interpretation of screening results. Additionally, they cover background information, such as proper referencing and disclosures of conflicts of interest. The first research question (RQ) directly asks: Does the health risk communication provided by LLMs reflect the best available evidence? The hypothesis is that LLMs largely fail to meet standard criteria for evidence-based health risk communication, as indicated by more than 50% of responses deviating, on average, from these guidelines. Using both MAPPinfo, an established assessment instrument, 39 and ebmNucleus, an assessment proposal derived from the Guideline Evidence-Based Health Communication, 3 Study 1 systematically examines the quality of outputs generated in response to repeated prompts concerning mammography for breast cancer screening (BC) and the PSA test for prostate cancer screening (PC). Prompt engineering refers to the effective interaction with LLMs to obtain optimal responses. The effect of prompt quality on LLM-generated responses has been demonstrated, for instance, in a March 2023 study on cardiovascular health advice, which used a non-validated expert-rating system to assess correctness, including the dimension of simplification in presentation. 40 However, it remains unknown whether prompt engineering also leads to more evidence-based responses for health decision-making. Therefore, Study 1 explores whether more informed prompts containing context knowledge increase LLM response quality. Accordingly, our RQ2 asks: Does a systematic variation of the contextual informedness of prompting increase the proportion of evidence-based LLM responses? We assume at least a moderately strong association between the grade of prompting informedness and the extent to which LLM responses are scored as evidence-based. We expect that our findings from this systematic variation of LLM prompting will align with prior studies that have used different accuracy benchmarks. Building on this, we aim to provide causal evidence that supporting laypeople in prompting can improve the quality of LLM responses they receive. However, since little is known about how effectively laypeople prompt without guidance, Study 2 first describes their LLM performance by assessing the compliance of LLM outputs with evidence-based standards—using the same instruments and conditions as in Study 1. Second, earlier studies have shown that prompt engineering training can improve response quality, as demonstrated, for example, with journalists. 41 . However, training programs face hurdles in practical implementation, such as time requirements, costs, and the need for human involvement. Beyond traditional education, boosting 42 could be a more feasible approach. Boosts–interventions that modify human cognition or the environment using behavioral insights to enhance people’s competences–have been proposed for navigating digital information architectures, including applications in fake news detection 43 and health information search. 44 Study 2 examines RQ3: Does a minimal boosting intervention that encourages laypeople to provide more informed prompts increase the proportion of evidence-based responses? To sum up, we aim to describe the limitations of LLM with regard to the standards of evidence-based medicine, demonstrate the potential for improvement through contextually more informed prompting, confirm the expected shortcomings of human prompting, and provide evidence that a simple intervention can help laypeople obtain more evidence-based health information. A study protocol was published in advance. 45 Results Study 1 LLMs’ responses had a median length of 211 (IQR: 158–278) words per prompt. All LLMs provided a response to every prompt. Our systematic variation of prompt informedness revealed the expected pattern across the tested LLMs (Figs. 1A-C), demonstrating that more informed prompting led to higher scores according to both assessment schemes MAPPinfo ( F (2,351) = 40·25, p < ·001, η p 2 = ·19) and ebmNucleus ( F (2,351) = 528·41, p < .001, η p 2 = ·75). The effect of informedness varied depending on the LLM as indicated by interactions for MAPPinfo ( F (4,351) = 3·48, p = ·008, η p 2 = ·04) and ebmNucleus ( F (4,351) = 12·11, p < ·001, η p 2 = ·12). Study 2 The sample (Table 3 ) consisted of 52·3% female, without any non-binary participants. Age ranged from 19 to 78 years, with a mean age of 46·4 years (SD = 15·5). The majority of participants were highly educated: 16·3% held a Master’s degree, 43·3% a Bachelor’s degree, and 4·3% a PhD. Additionally, 20·7% had completed school education up to the age of 18, while 9·7% had obtained professional or technical qualifications. A small proportion (5·7%) reported having no formal education beyond the age of 16. A total of 63·0% of participants reported having at least some experience with LLMs, and notably, 31·7% used them at least once per month to seek health information. On average, participants submitted 3.0 prompts (SD = 1·0). Those prompts had a median length of seven words (IQR: 5–9). All LLMs generated a response for every prompt, and the intervention increased the prompt word count ( p = ·032). The LLMs' responses to these prompts had a median length of 188 words (IQR: 126–254) per response. Independent of the specific LLM, the information elicited in the control condition reached only approximately 17% of the possible maximum score in MAPPinfo and 13% in ebmNucleus. A regression analysis of the control condition ( F (6, 144) = 7·75, p < ·001) revealed that participants generated higher quality information (ebmNucleus) when they had a higher level of education ( β standardized = ·16, p = ·032) and more experience with LLMs ( β = ·26, p = ·007).However, more frequent LLM usage for seeking health information was negatively associated with information quality ( β = -·25, p = ·009). Neither age ( p = ·219) nor gender ( p = ·952) had a significant influence, when controlling for the chosen topic, with prostate cancer screening prompts yielding higher-quality information ( β standardized = ·41, p = ·003). The boosting intervention for more informed prompting moderately improved the quality of elicited health information across all LLMs (Figs. 2A-C), as indicated by ANOVA main effects for MAPPinfo ( F (1,294) = 15·26, p < ·001, η p 2 = .05) and ebmNucleus ( F (1,294) = 12·12, p = ·001, η p 2 = ·04). The effect of the intervention neither varied for MAPPinfo ( F (2,294) = 0·32, p = ·726) nor for ebmNucleus ( F (1,294) = 1·18, p = ·307) across the LLMs. Discussion Our findings highlight several key insights into the quality of health information provided by LLMs and the role of prompt-informedness in improving response quality. First, our findings reinforce previous concerns that LLMs are not designed to inform health decision making and, in particular, they do not comply with the standards of evidence-based medicine. Even when well-informed prompts were crafted by researchers—ensuring maximum accordance with information requirements—not even 50% of maximum compliance could be reached with ChatGPT and Gemini responses (somewhat better in Le Chat). Second, our results confirm that the quality of elicited information directly depends on prompt informedness. This was systematically demonstrated by the gradual prompt manipulation, which aligned with the information requirements for informed decision-making. Importantly, this relationship was observed across both assessment tools—ebmNucleus, which explicitly focuses on these information requirements, and MAPPinfo, an established but broader quality assessment instrument. Third and most notably, we provide causal evidence that a simple boosting intervention—reminding users to inquire about possible consequences of medical options—improved the quality of LLM-generated health information. This extends previous research on digital navigation boosts, 42 , 44 demonstrating that those behavioral interventions can help users seek better information. Moreover, our boost could be easily implemented at LLM interfaces. Despite the strengths of our study, several limitations warrant consideration. The relatively low scores obtained with MAPPinfo may suggest that traditional health information assessment tools need to be carefully adapted for evaluating LLM outputs. However, this does not imply that the fundamental quality criteria for evidence-based health information should change. These criteria are derived from established evidence and codified in guidelines, forming the basis of our investigation. MAPPinfo was originally designed for structured health information sources, such as websites and brochures. While LLM-generated responses are dynamic and interactive, they should still adhere to the same evidence-based standards. If LLM-generated health information consistently fails to meet these standards, this raises concerns about their appropriateness as a source for informed medical decision-making rather than the validity of the assessment tools themselves. The lower scores observed in participant-driven prompting with ebmNucleus may be explained by the task instructions. Participants were asked to seek general information about BC or PC screening, rather than specifically about mammography or PSA testing. This methodological choice influenced scoring outcomes because broad inquiries (e.g., "Tell me about breast cancer screening") may generate responses covering multiple screening technologies, leading to less targeted benefit-harm presentations. While this directly affected only 8 out of 38 points in MAPPinfo, it influenced 18 out of 27 points in ebmNucleus, where a structured benefit-risk assessment is central. Nevertheless, these findings reflect real-world search behaviors, where users may not always phrase their questions in the most targeted manner. Another limitation is that our Study 1 relied on single-turn prompts, rather than engaging in multi-turn conversations. This should ensure controlled and systematic prompt variations, making it easier to isolate the effects of prompt informedness. Given that conversational interactions with LLMs can improve information quality by allowing users to clarify, elaborate, or challenge information, future studies should explore the effect of conversation. However, the LLM conversations in Study 2 did not lead to responses of higher quality. Furthermore, it is possible that today’s LLMs may yield better evidence-based responses. However, there is currently no indication that LLMs are explicitly trained on evidence-based health communication principles. Additionally, while some newer models have real-time web access, this does not necessarily imply better compliance with EBM. Future studies should assess how real-time retrieval mechanisms influence response accuracy and bias in health-related LLM interactions. Our study suggests that differences in education level and prior LLM experience influence the quality of information users retrieve. While more experienced users generated higher-quality responses, those who frequently used LLMs for health-related searches retrieved lower-quality information. This highlights a potential risk of digital inequality, where less-experienced users may unknowingly rely on incomplete or misleading AI-generated health information. Future research should investigate whether LLM interfaces could be adapted to mitigate such disparities, for example, by providing adaptive guidance based on user expertise levels. LLMs generate output based on probabilistic language modeling, meaning they can reinforce existing biases in health communication. For instance, certain populations may receive different risk-benefit framing depending on how they phrase their queries. Given that health decisions disproportionately affect marginalized groups, fairness considerations in AI-generated health information should be systematically addressed. One possible approach is to audit LLM-generated health responses across demographic groups to detect and correct disparities. To sum up, LLMs are used as health information sources, yet our findings emphasize that they do not inform appropriately. While informedness plays a crucial role in improving response quality, even very informed prompts fail to ensure full compliance with EBM standards. As LLMs continue to evolve, AI developers must focus not only on integrating the best available clinical evidence but also on leveraging behavioral science insights to optimize how information can be summarized and presented in balanced and transparent way. 46 Then can AI-driven health communication support knowledge acquisition, correct risk perceptions, and enable informed decision-making. But even in the light of further improvement, laypeople must also be informed about how well LLMs perform and where these models fall short, as users may hold unrealistic expectations about algorithm quality. 47 Finally, as it is true for every algorithm system that is to be widely implemented in healthcare, the system’s stakeholders need to understand related consequences for different groups and stakeholders 48 ; ideally with the help of large randomized trials. Methods Study 1 Study 1 employed systematic prompting variations to assess the compliance of LLM-generated health risk information with evidence-based communication standards. Reporting followed the STROBE 49 and TRIPOD-LLM 50 guidelines. Design and Data collection This content analysis aimed to systematically evaluate the quality of responses generated by three LLMs—OpenAI’s ChatGPT (gpt-3·5-turbo), Google Gemini (1·5-Flash), and Mistral AI Le Chat (mistral-large-2402)—when prompted with health-related queries. These LLMs where selected based on their accessibility and public awareness in Germany during the summer 2024, as they were available without payment or the need to disclose personal data (except for an email address). At that time, specialized medical LLMs did not meet these criteria. The study followed a pre-registered content analysis design (AsPredicted, Registration No. 180732), and was conducted between June and July 2024. Since no human participants were involved and no personal data were used, ethical review was not required. Our systematic prompting strategy (Table 1) included 18 prompts (Tables S1-2) for both mammography and for the PSA test. The strategy varied prompts across three levels of contextual informedness: (1) highly informed prompts, which included technical keywords, (2) moderately informed prompts, which used more generic synonyms, and (3) low-informed prompts, which omitted key terms entirely. Additionally, the strategy incorporated six crucial information requirements necessary for supporting informed medical decision-making–one of the key principles of evidence-based medicine. These requirements enable individuals to weigh the potential benefits and harms of undergoing versus not undergoing a certain medical procedure, while also considering underlying uncertainties and the screening test-specific conditional probabilities. LLMs are probabilistic models that produce variable outputs even when given the exact same prompt. To ensure a valid assessment—while many previous studies have relied on insufficient prompt repetitions (e.g., around 1 to 4) 11,28 —we designed our study with a sample size of twenty independent trials per prompt to account for low probability responses. The three LLMs were accessed via API, without output length limitations, resulting in a total of 2,160 prompts. Mammography for BC and the PSA test for PC screening were selected as study topics because they represent two of the most common cancers worldwide, affecting millions of individuals and posing significant health implications for screening decisions. Both involve complex risk-benefit considerations in screening and treatment, require clear, evidence-based communication to support informed decision-making. Additionally, BC and PC screenings have been subject to distorted public perceptions, 51,52 as well as ongoing debate due to concerns over overdiagnosis, false positives, and varying guidelines. These factors make them ideal case studies for assessing whether LLMs can provide accurate, guideline-adherent health information. Scoring and analysis of LLM responses Two independent researchers and research assistants who were blinded to each other’s assessments, the specific LLM, the prompt, and its level of informedness scored the LLM responses using a validated scoring instrument, MAPPinfo 39 for evidence-based health information quality, and, in addition to that, the proposal ebmNucleus, which draws on the guideline for evidence-based health information 3 (Table S3). Both tools share certain criteria (Table 2) but differ in their focus: while MAPPinfo provides a broad assessment of information quality, ebmNucleus concentrates on the core aspects of health decision-making, minimizing the influence of contextual factors irrelevant to most LLMs (e.g., authorship, graphical elements). Each LLM response was systematically labeled and scored accordingly. Here, the evidence underlying the content was assessed according to the best available evidence (summarized under www.hardingcenter.de/en/transfer-and-impact/fact-boxes) from 2022, aligning with the training corpora of the LLMs used. While responses were consistently scored on a zero to two-point scale in MAPPinfo, responses concerning patient-relevant benefits and harms, referencing of single event probabilities, benefit and harm numbers, evidence quality, and interpretation of screening results was scored on a zero to three-point scale in ebmNucleus, while all other aspects were scored on a binary zero to one-point scale. Responses classified as ‘hallucinations’ or absurd statements were scored accordingly. The raters were research assistants (RAs) and experienced researchers who had previously conducted similar coding tasks in multiple studies. The latter had expertise in evidence-based health communication. Each rater was instructed how to apply the scoring instruments. Two RA’s (each 50%) and one researcher independently scored LLM responses according to MAPPinfo; dyads from a pool of five RA’s independently scored according to ebmNucleus (Table S3). The highest possible scoring (three points) of a test result interpretation required, for instance, providing the probability that a positive result may be due to underlying disease (=positive predictive value). The lowest possible score was applied when there was no hint on the limits of the screening test or just one-sided (false negatives). Interrater reliability across all LLM outputs, calculated using Cohen’s kappa coefficient, was ϰ = ·41 for MAPPinfo and ϰ = ·38 for ebmNucleus. Because the reliability was too low, a third rater, a researcher (FGR and CW, respectively), who was blinded to the previous ratings, rated cases of disagreement. Their experienced ratings resulted each in ϰ = ·94 agreement with the respective prior raters. Raw sum scores were calculated for both MAPPinfo and ebmNucleus. Average sum scores were calculated for each LLM and prompt category and converted into the rate achieved regarding perfect evidence-based criteria. Descriptive statistics and variance analyses were conducted to examine the effect of prompting informedness on scoring outcomes. Study 2 Study 2 employed systematic investigations of laypeople’s prompting to assess the extent to which LLM-generated responses were evidence-based and how this could be improved through a boosting intervention. Reporting followed the SROBE 49 and TRIPOD-LLM 50 guidelines. Sample To detect a moderate ANOVA main effect (partial eta squared = ·06) when comparing two between-subjects conditions (with and without intervention), we required a minimum sample size of n = 237 participants. To approximate simplified census data of Great Britain in terms of sex, age, and ethnicity, we recruited n = 300 adult participants from an online-representative pool of participants via Prolific. Their remuneration was €1·30 upon completion of the study. Design We used a 2x3 between-subjects design: participants were randomly assigned either to standard prompting instructions (control) or enhanced prompting instructions (boosting intervention) and to one of three LLMs; OpenAI ChatGPT (gpt-3·5-turbo), Google Gemini (1·5-Flash), or Mistral AI Le Chat (mistral-large-2402). To ensure that participants could generate detailed prompts and receive more comprehensive responses, the survey was restricted to tablets and computers. Typing on smartphones is not equivalent to using a full keyboard, which could affect both the quality and complexity of the prompts as well as the resulting responses. A minimal boosting intervention encouraged participants to consider the possible consequences of their choices as follows: “Please consider the OARS rule: You need to know your Options, the Advantages and Risks of each, and how Steady they are to happen.” Participants in the control group did not receive this intervention. To ensure a balanced distribution of participants across study groups, we employed block randomization in SoSci Survey where we hosted the study. Both participants and researchers were blinded to the assigned LLM and the type of prompting instructions to prevent bias in interaction and response evaluation. The study was pre-registered with the German Clinical Trials Register (DRKS) (Registration No. DRKS00035228) and received ethical approval from the Ethics Committee of the University of Potsdam (Approval No. 52/2024). It was conducted in October 2024. Pretest A pre-test that involved n = 20 participants was conducted to verify the effectiveness of the randomization procedures, test the technical functionality of the survey platform and its integration with the LLM APIs. Additionally, it measured the average completion time and identified potential issues related to participant burden. User feedback was collected to address any usability concerns, and necessary adjustments were made to optimize the study protocol for the main trial. The pre-test also assessed the study materials, particularly the clarity and comprehensibility of the survey questions and instructions. Procedure and material Participants received a brief introduction outlining the study’s purpose, which involved interacting with one of three LLMs, all preset as a “helpful assistant” to obtain health information about either BC or PC screening. After giving informed consent, participants completed the questionnaire, providing basic demographic details such as gender, age, and education level. Participants could choose between BC and PC screening, regardless of their self-reported gender. This approach aimed to enhance engagement and data relevance while maintaining the integrity of the study design. They were then instructed to gather information on their chosen topic, which was expected to increase variance in information quality compared to focusing solely on mammography and PSA testing. Participants interacted with the chatbot by entering their queries (prompts) via the LLM’s API, and the generated responses were collected alongside the prompts for systematic analysis. To ensure comparability across participants and maintain a controlled interaction length, each participant could submit a maximum of four prompts. At the end of the session, participants reported their frequency of LLM usage, prior experience with such models, and their attitude on shared decision-making. On average, participants completed the survey (Table S4) in five minutes. Measures In addition to demographic measures such as self-reported gender, age, and education level, we assessed LLM usage frequency, prior experience, and preferred approach to medical decision-making using specific items. LLM usage frequency was measured with the question, “How often have you used computer programs like ChatGPT (large language models) to get health information?” Response options included: Never, About 1–5 times per year, About 1–2 times per month, About 1–2 times per week, More frequently. Prior general experience was assessed with the item, “Please rate your experience with computer programs like ChatGPT (large language models).” Participants responded on a scale: Definitely no experience, Rather few experience, Some experience, Rather much experience, Definitely much experience. Their attitude on shared decision-making was elicited with statements about decision responsibility between GP and oneself. Scoring and analysis To evaluate the responses generated by the LLMs, scoring followed the same procedure as in Study 1. Two raters independently (one RA, one researcher according to MAPPinfo; two researchers according to ebmNucleus) scored the 300 LLM outputs. Interrater reliability, calculated using Cohen’s kappa coefficient, was ϰ = ·44 for MAPPinfo and ϰ = ·71 for ebmNucleus. A third rater, a researcher (FGR or CW), who was blinded to the previous ratings, scored cases of disagreement. Their experienced ratings resulted in ϰ = ·99 (MAPPinfo) and ϰ = ·98 (ebmNucleus) agreement with the respective prior raters. Raw sum scores were calculated for both MAPPinfo and ebmNucleus. The average sum scores were computed for each LLM and prompt category and converted into the rate of adherence to perfect evidence-based criteria. Descriptive statistics were calculated for demographic variables such as self-reported gender, age, and education, as well as for LLM usage frequency and prior experience. Variance analyses were conducted to examine the impact of prompt informedness on response quality, while interaction effects between LLM type and prompting instructions were also evaluated. Regression analyses explored the influence of demographic characteristics, LLM usage and prior experience. All statistical tests were conducted at a significance level of p < ·05, and effect sizes were calculated to determine the practical significance of findings. Declarations Data availability All data supporting the findings of this study are available in the Supplementary Section of this article. Acknowledgements This study did not receive any funding. Author contributions FGR was responsible for the conceptualization and development of the study methodology, conducted the statistical analysis, interpreted the results, and co-drafted the original manuscript. AC and SKP worked on the conceptual planning, the development and structuring of the study, and oversaw the literature. MZ was responsible for the technical planning and implementation. AS contributed to the study design, validated the methodology, and helped revise the manuscript. CW was responsible for questionnaire development and pretesting, supervised the study, and managed project administration. He ensured the rigorous execution of the study, contributed to the conceptualization, study design, methodology, and data collection, co-drafted the original manuscript, and was responsible for substantial manuscript revisions and editing. He also serves as the primary contact and guarantor of the manuscript. All authors had full access to the underlying data reported in the manuscript and verified its accuracy. They collectively contributed to the interpretation of results and were involved in manuscript writing. Competing Interests All authors declare no financial or non-financial competing interests. References Calixte, R., Rivera, A., Oridota, O., Beauchamp, W. & Camacho-Rivera, M. Social and demographic patterns of health-related internet use among adults in the United States: a secondary data analysis of the health information national trends survey. International Journal of Environmental Research and Public Health 17 , 6856 (2020). Riera, R. et al. Strategies for communicating scientific evidence on healthcare to managers and the population: a scoping review. Health Research Policy and Systems 21 , 71, doi:10.1186/s12961-023-01017-2 (2023). Lühnen, J., Albrecht, M., Mühlhauser, I. & Steckelberg, A. Leitlinie evidenzbasierte Gesundheitsinformation. (EBM network, 2017). Lühnen, J., Albrecht, M., Hanßen, K., Hildebrandt, J. & Steckelberg, A. Leitlinie evidenzbasierte Gesundheitsinformation: Einblick in die Methodik der Entwicklung und Implementierung. Zeitschrift für Evidenz, Fortbildung und Qualität im Gesundheitswesen 109 , 159-165 (2015). Johnson, S. B. et al. Cancer misinformation and harmful information on Facebook and other social media: a brief report. JNCI: Journal of the National Cancer Institute 114 , 1036-1039 (2022). Han, T. et al. MedAlpaca--an open-source collection of medical conversational AI models and training data. arXiv preprint arXiv:2304.08247 (2023). Park, Y.-J. et al. Assessing the research landscape and clinical utility of large language models: a scoping review. BMC Medical Informatics and Decision Making 24 , 72, doi:10.1186/s12911-024-02459-6 (2024). Choudhury, A. & Shamszare, H. Investigating the impact of user trust on the adoption and use of ChatGPT: survey analysis. Journal of Medical Internet Research 25 , e47184 (2023). Mendel, T., Singh, N., Mann, D. M., Wiesenfeld, B. & Nov, O. Laypeople’s Use of and Attitudes Toward Large Language Models and Search Engines for Health Queries: Survey Study. Journal of Medical Internet Research 27 , e64290 (2025). Anastasio, A. T., Mills IV, F. B., Karavan Jr, M. P. & Adams Jr, S. B. Evaluating the quality and usability of artificial intelligence–generated responses to common patient questions in foot and ankle surgery. Foot & Ankle Orthopaedics 8 , 24730114231209919 (2023). Wang, G. et al. AI's deep dive into complex pediatric inguinal hernia issues: a challenge to traditional guidelines? Hernia 27 , 1587-1599 (2023). Bedi, S. et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA 333 , 319-328, doi:10.1001/jama.2024.21700 (2025). Onder, C., Koc, G., Gokbulut, P., Taskaldiran, I. & Kuskonmaz, S. Evaluation of the reliability and readability of ChatGPT-4 responses regarding hypothyroidism during pregnancy. Scientific reports 14 , 243 (2024). Amin, K., Doshi, R. & Forman, H. P. in Healthcare. 100731 (Elsevier). Haver, H. L. et al. Evaluating the use of ChatGPT to accurately simplify patient-centered information about breast cancer prevention and screening. Radiology: Imaging Cancer 6 , e230086 (2024). Peng, C. et al. A study of generative large language model for medical research and healthcare. NPJ digital medicine 6 , 210 (2023). Davis, R. J. et al. Evaluation of oropharyngeal cancer information from revolutionary artificial intelligence chatbot. The Laryngoscope 134 , 2252-2257 (2024). Razdan, S., Siegal, A. R., Brewer, Y., Sljivich, M. & Valenzuela, R. J. Assessing ChatGPT’s ability to answer questions pertaining to erectile dysfunction: can our patients trust it? International Journal of Impotence Research 36 , 734-740 (2024). Sarraju, A. et al. Appropriateness of cardiovascular disease prevention recommendations obtained from a popular online chat-based artificial intelligence model. Jama 329 , 842-844 (2023). Haver, H. L. et al. Appropriateness of breast cancer prevention and screening recommendations provided by ChatGPT. Radiology 307 , e230424 (2023). Tailor, P. D. et al. Appropriateness of ophthalmology recommendations from an online chat-based artificial intelligence model. Mayo Clinic Proceedings: Digital Health 2 , 119-128 (2024). Alapati, R. et al. Evaluating insomnia queries from an artificial intelligence chatbot for patient education. Journal of Clinical Sleep Medicine 20 , 583-594 (2024). Biswas, S., Logan, N. S., Davies, L. N., Sheppard, A. L. & Wolffsohn, J. S. Assessing the utility of ChatGPT as an artificial intelligence‐based large language model for information to answer questions on myopia. Ophthalmic and Physiological Optics 43 , 1562-1570 (2023). Ayers, J. W. et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine 183 , 589-596 (2023). Caglar, U. et al. Evaluating the performance of ChatGPT in answering questions related to benign prostate hyperplasia and prostate cancer. Minerva Urology and Nephrology 75 , 729-733 (2023). Liu, J., Zheng, J., Cai, X., Wu, D. & Yin, C. A descriptive study based on the comparison of ChatGPT and evidence-based neurosurgeons. Iscience 26 (2023). Yan, S. et al. Assessment of the Reliability and Clinical Applicability of ChatGPT’s Responses to Patients’ Common Queries About Rosacea. Patient preference and adherence , 249-253 (2024). Wang, L. et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. npj Digital Medicine 7 , 41 (2024). Johnson, S. B. et al. Using ChatGPT to evaluate cancer myths and misconceptions: artificial intelligence and cancer information. JNCI cancer spectrum 7 , pkad015 (2023). Alessandri-Bonetti, M., Giorgino, R., Naegeli, M., Liu, H. Y. & Egro, F. M. Assessing the soft tissue infection expertise of ChatGPT and Bard compared to IDSA recommendations. Annals of Biomedical Engineering 52 , 1551-1553 (2024). Huo, B. et al. Dr. GPT will see you now: the ability of large language model-linked chatbots to provide colorectal cancer screening recommendations. Health and Technology 14 , 463-469 (2024). Cappellani, F., Card, K. R., Shields, C. L., Pulido, J. S. & Haller, J. A. Reliability and accuracy of artificial intelligence ChatGPT in providing information on ophthalmic diseases and management to patients. Eye 38 , 1368-1373 (2024). Cheong, R. C. T. et al. Artificial intelligence chatbots as sources of patient education material for obstructive sleep apnoea: ChatGPT versus Google Bard. European Archives of Oto-Rhino-Laryngology 281 , 985-993 (2024). Mejia, M. R. et al. Use of ChatGPT for determining clinical and surgical treatment of lumbar disc herniation with radiculopathy: a North American Spine Society guideline comparison. Neurospine 21 , 149 (2024). Wang, G. et al. Potential and limitations of ChatGPT 3.5 and 4.0 as a source of COVID-19 information: comprehensive comparative analysis of generative and authoritative information. Journal of Medical Internet Research 25 , e49771 (2023). Working group GPGI. Gute Praxis Gesundheitsinformation [Good practice health information]. Zeitschrift für Evidenz, Fortbildung und Qualität im Gesundheitswesen 110 , 85-92, doi:10.1016/j.zefq.2015.11.005 (2016). Pan, A., Musheyev, D., Bockelman, D., Loeb, S. & Kabarriti, A. E. Assessment of artificial intelligence chatbot responses to top searched queries about cancer. JAMA oncology 9 , 1437-1440 (2023). Walker, H. L. et al. Reliability of medical information provided by ChatGPT: assessment against clinical guidelines and patient information quality instrument. Journal of Medical Internet Research 25 , e47479 (2023). Kasper, J. et al. MAPPinfo‐mapping quality of health information: Validation study of an assessment instrument. PloS one 18 , e0290027 (2023). Lautrup, A. D. et al. Heart-to-heart with ChatGPT: the impact of patients consulting AI for cardiovascular health advice. Open heart 10 , e002455 (2023). Bashardoust, A., Feng, Y., Geissler, D., Feuerriegel, S. & Shrestha, Y. R. The Effect of Education in Prompt Engineering: Evidence from Journalists. arXiv preprint arXiv:2409.12320 (2024). Herzog, S. M. & Hertwig, R. Boosting: Empowering citizens with behavioral science. Annual Review of Psychology 76 (2025). Kozyreva, A. et al. Toolbox of individual-level interventions against online misinformation. Nature Human Behaviour , 1-9 (2024). Rebitschek, F. G. & Gigerenzer, G. Einschätzung der Qualität digitaler Gesundheitsangebote: Wie können informierte Entscheidungen gefördert werden? [Assessing the quality of digital health services: How can informed decisions be promoted?]. Bundesgesundheitsblatt-Gesundheitsforschung-Gesundheitsschutz 63 , 665-673, doi:10.1007/s00103-020-03146-3 (2020). Rebitschek, F. G. W., Christoph. Study Protocol for a Two-Phase Randomized Evaluation of Large Language Models in Adherence to Evidence-Based Health Communication Guidelines for Breast and Prostate Cancer Screening: The Role of User Prompt Specificity and Minimal Interventions (BOOST-AI). Version 1.0 from October 16th 2024. . (2024). McDowell, M., Rebitschek, F. G., Gigerenzer, G. & Wegwarth, O. A simple tool for communicating the benefits and harms of health interventions. MDM Policy & Practice 1 , 2381468316665365 (2016). Rebitschek, F. G., Gigerenzer, G. & Wagner, G. G. People underestimate the errors made by algorithms for credit scoring and recidivism prediction but accept even fewer errors. Scientific Reports, 11 , doi:10.1038/s41598-021-99802-y (2021). Wilhelm, C., Steckelberg, A. & Rebitschek, F. G. Benefits and harms associated with the use of AI-related algorithmic decision-making systems by healthcare professionals: a systematic review. The Lancet Regional Health–Europe 48 (2025). von Elm, E. et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. J Clin Epidemiol 61 , 344-349, doi:10.1016/j.jclinepi.2007.11.008 (2008). Gallifant, J. et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med 31 , 60-69, doi:10.1038/s41591-024-03425-5 (2025). Wegwarth, O. et al. What do European women know about their female cancer risks and cancer screening? A cross-sectional online intervention survey in 5 European countries. BMJ Open 8 , doi:doi:10.1136/bmjopen-2018-023789 (2018). Gigerenzer, G., Mata, J. & Frank, R. Public knowledge of benefits of breast and prostate cancer screening in Europe. Journal of the National Cancer Institute 101 , 1216-1220 (2009). Tables Table 1. Prompting strategy underlying systematic testing Information requirements Low informed Moderately informed Highly informed Listing possible benefits and harms Asking for guidance Asking for consequences Asking for benefits and harms Explaining single event probabilities Asking for the chance Asking for an explanation Asking for reference conditions Quantifying screening benefits Asking for the advantage Asking for a probability Asking for an absolute effect Quantifying numerical screening harms Asking for the disadvantage Asking for a probability Asking for an absolute effect Informing about evidence quality Asking for an estimate Asking for reliability Asking for the study quality Interpreting test result Asking for positive result meaning Asking for the positive predictive value Providing input for conditional reasoning Table 2. Criteria of the health information quality scoring instruments used in Study 2. MAPPinfo criteria Overlapping criteria ebmNucleus criteria “Informed decision” Benefit numbers Conflict of interest (COI) Diagnostic quality Evidence methodology Final update Harm numbers No narratives References Stochastic uncertainty Authors Framing - Health problem COI management Natural course/prevalence Options Recipients’ definition Suitable graphics Evidence quality Neutrality Patient-relevant benefits and harms Referencing single event probabilities Referral to support Table 3. Sample description according to intervention condition. Intervention Control Difference (p) Gender (% female) 53·7 51·0 - Age in years (M[SD]) 46·5 [15·8] 46·3 [15·3] - Education - PhD (%) 5·4 3·3 Master’s degree (%) 18·1 14·6 Bachelor’s degree (%) 39·6 47·0 School education up to age 18 (%) 18·8 22·5 Professional or technical qualifications (%) 12·1 7·3 No formal education (%) 6·0 5·3 At least some experience with LLM (%) 65·8 60·3 - At least once per month LLM health info. (%) 30·2 33·1 - Preferred topic - Breast cancer screening (%) 55·0 52·3 Prostate cancer screening (%) 45·0 47·7 Number of prompts out of four (M[SD]) 3·0 [1·0] 2·9 [1·0] - Prompt word count (M[SD]) 8·2 [6·4] 7·3 [3·4] ·032 Additional Declarations No competing interests reported. Supplementary Files TableS.docx Cite Share Download PDF Status: Published Journal Publication published 09 Jun, 2025 Read the published version in npj Digital Medicine → Version 1 posted Editorial decision: Revision requested 30 Apr, 2025 Reviews received at journal 27 Apr, 2025 Reviewers agreed at journal 15 Apr, 2025 Reviews received at journal 03 Apr, 2025 Reviewers agreed at journal 01 Apr, 2025 Reviewers agreed at journal 31 Mar, 2025 Reviewers invited by journal 25 Mar, 2025 Editor assigned by journal 13 Mar, 2025 Submission checks completed at journal 13 Mar, 2025 First submitted to journal 13 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6220209","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":428584185,"identity":"4a737fbc-ff18-426a-a0bc-b497afffa2b2","order_by":0,"name":"Felix G. Rebitschek","email":"","orcid":"","institution":"University of Potsdam","correspondingAuthor":false,"prefix":"","firstName":"Felix","middleName":"G.","lastName":"Rebitschek","suffix":""},{"id":428584186,"identity":"83089881-c8ed-4869-92d3-53d5072bde78","order_by":1,"name":"Alessandra Carella","email":"","orcid":"","institution":"University of Padova","correspondingAuthor":false,"prefix":"","firstName":"Alessandra","middleName":"","lastName":"Carella","suffix":""},{"id":428584187,"identity":"6b8d5c5a-cd0d-49c3-8382-eb4730f714aa","order_by":2,"name":"Silja Kohlrausch-Pazin","email":"","orcid":"","institution":"University of Potsdam","correspondingAuthor":false,"prefix":"","firstName":"Silja","middleName":"","lastName":"Kohlrausch-Pazin","suffix":""},{"id":428584188,"identity":"d78d0649-e56a-4ceb-8307-46f936e22b5b","order_by":3,"name":"Michael Zitzmann","email":"","orcid":"","institution":"University of Potsdam","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Zitzmann","suffix":""},{"id":428584189,"identity":"442393be-dc6e-4cac-ae04-8a1555243980","order_by":4,"name":"Anke Steckelberg","email":"","orcid":"","institution":"Martin Luther University Halle-Wittenberg","correspondingAuthor":false,"prefix":"","firstName":"Anke","middleName":"","lastName":"Steckelberg","suffix":""},{"id":428584190,"identity":"f338340b-14c1-449f-beb8-5a91929a8487","order_by":5,"name":"Christoph Wilhelm","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABOklEQVRIie3SMUvDQBQH8AcdujzJeqLETyBcCVyRBP0qdxSuS+gSKG5myqS4Kn6JiiCOV27okg8QaYd2cXJonVIE9dKmFpLW2eH+cBzk+PHeuxyAjc1/DAFQxU7NUrMYjqEZA/D1YfGxsZvwkogYEFCtCf5BYENgRUhZYh9xHm5mw2Ue9NqALSVeAjy9f2fTKUzcC0c/fmTgu9UikxHVyGV0FiNVIpXIxmGbcnjzkMjoKISuVyGUSNDAtRiog4ESiS4II+JLi2uC1BBtuq2RYb4l38heU2bGMcRJvU9DrnYQhVuikGVYEghZUYVXZ8lMYyhlRLUzN6SDLJV9spolk30/pN1WpYpzJxuLPAh6dJR05ovk3GUj/XyYmxtr3uqncXjpn9R+zDq8ev9q0/ke8PtC6sTGxsbGBn4AHAl4MItxRJAAAAAASUVORK5CYII=","orcid":"","institution":"International Graduate Academy (InGrA), Martin Luther University Halle-Wittenberg","correspondingAuthor":true,"prefix":"","firstName":"Christoph","middleName":"","lastName":"Wilhelm","suffix":""}],"badges":[],"createdAt":"2025-03-13 12:53:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6220209/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6220209/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41746-025-01752-6","type":"published","date":"2025-06-09T15:57:09+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":78888268,"identity":"75029d78-fe56-4b49-849d-3b6944dd1df2","added_by":"auto","created_at":"2025-03-20 09:59:01","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":72326,"visible":true,"origin":"","legend":"\u003cp\u003eA-C. The dependence of information quality as elicited with the help of two metrics (MAPPinfo and ebmNucleus) on the informedness of prompts is demonstrated across Gemini (A), Le Chat (B), and ChatGPT (C). Error bars show the standard error of the average rate of fulfilled criteria across trials.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6220209/v1/3f8692d2996f0d4f53ba9892.png"},{"id":78888664,"identity":"9e405e6b-2a86-47f8-a59e-1a95a0bc94ac","added_by":"auto","created_at":"2025-03-20 10:07:00","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":70564,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA-C. \u003c/strong\u003eThe improvement of information quality as elicited with the help of two metrics (MAPPinfo and ebmNucleus) with the help of a boost is demonstrated across Gemini (A), Le Chat (B), and ChatGPT (C). Error bars show the standard error of the average rate of fulfilled criteria across participants.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6220209/v1/6eba2bb8b770dff17fd3550e.png"},{"id":84726463,"identity":"bcabcf8e-fcec-4108-b6d6-4d47bb6f02c1","added_by":"auto","created_at":"2025-06-16 16:04:55","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":882656,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6220209/v1/a038228e-b7ab-42c3-b36d-9a1c2c2a2c72.pdf"},{"id":78888266,"identity":"a0315b0b-dc19-4687-a712-836da7079f49","added_by":"auto","created_at":"2025-03-20 09:59:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":23694,"visible":true,"origin":"","legend":"","description":"","filename":"TableS.docx","url":"https://assets-eu.researchsquare.com/files/rs-6220209/v1/0b0797dbdd601fe8df4a1293.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluating Evidence-Based Communication through Generative AI using a Cross-Sectional Study with Laypeople Seeking Screening Information","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe internet has become a primary source of health information for people in Western countries.\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e However, many online resources fail to adhere to evidence-based communication standards.\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e As the practical implementation of guidelines for evidence-based health communication\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e remains challenging, much of the health information available online does not support informed health decision-making, it lacks necessary accuracy and rigor, and in some cases even spreads misinformation, particularly in areas such as cancer.\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e The rapid emergence of artificial intelligence (AI) tools, particularly large language models (LLMs) such as OpenAI's ChatGPT, Mistral AI's Le Chat, has opened novel possibilities for digital health communication \u0026ndash; even in professional settings.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e Laypeople are increasingly turning to these platforms to seek answers to health-related questions.\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e However, can LLM users reliably obtain correct information that complies with established health communication guidelines? To address this, we conduct two complementary studies aimed at evaluating and enhancing the quality of health communication provided by LLMs.\u003c/p\u003e \u003cp\u003eAlthough LLMs are not designed for health communication, laypeople use them to gather health information. In a convenience online sample of adult panelists from February to March 2023, 7\u0026middot;2% reported regularly using LLMs for health-related topics\u0026mdash;often in combination with search engines and online health communities.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e This number increased significantly, with 73\u0026middot;6% of an online community convenience sample between December 2023 and January 2024 and 32\u0026middot;6% of a U.S. online panel in February 2024 using LLMs for health information.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThough often correct, LLM-generated responses to health-related queries show substantial variability.\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e The \u0026lsquo;accuracy\u0026rsquo; or \u0026lsquo;reliability\u0026rsquo; of LLM responses has been subject to numerous descriptive studies, including a recent systematic review of 88 patient education studies.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e Yet, what does accuracy or reliability truly mean in this context? LLM-generated responses to health-related questions has been assessed in various ways, such as readability,\u003csup\u003e\u003cspan additionalcitationids=\"CR14 CR15 CR16 CR17\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e expert-rated appropriateness of factual explanations,\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan additionalcitationids=\"CR19 CR20 CR21 CR22 CR23 CR24 CR25 CR26\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e and agreement with health authorities and guidelines.\u003csup\u003e\u003cspan additionalcitationids=\"CR29 CR30 CR31 CR32 CR33 CR34\" citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e However, what remains missing, is the use of the best available evidence as the ground truth.\u003c/p\u003e \u003cp\u003eLLM information and communication studies consistently lack a well-defined health-care premise (e.g. informed decision-making). This absence undermines both the criteria for response assessment (e.g. compliance with evidence-based communication standards) and the structured selection of prompts (e.g. aligning with information requirements). Moreover, these studies often fail to ground their assessments in the best available evidence, such as high-quality evidence synthesis. Evidence-based guidelines, such as those developed by the Working Group for Good Practice in Health Information (GPHI), provide a structured framework for health communication.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e They emphasize, for instance, the importance of using numerical data to present risks and benefits based on the best available evidence. Regarding validated instruments for assessing information quality, at least three LLM studies\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e have applied the DISCERN standards, which from today's perspective, are considered outdated, while one study utilized a validated patient information quality instrument.\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003e The primary aim of Study 1, therefore, is to assess the extent to which content generated by state-of-the-art LLMs\u0026mdash;specifically OpenAI\u0026rsquo;s ChatGPT, Google Gemini, and Mistral AI\u0026mdash;adheres to guideline-based criteria, with a particular focus on essential aspects of health risk communication for informed decision-making. These criteria encompass clear communication of benefits and harms, inclusion of reference class information, the appropriate presentation of numerical effects, the quality of evidence, and accurate interpretation of screening results. Additionally, they cover background information, such as proper referencing and disclosures of conflicts of interest. The first research question (RQ) directly asks: Does the health risk communication provided by LLMs reflect the best available evidence? The hypothesis is that LLMs largely fail to meet standard criteria for evidence-based health risk communication, as indicated by more than 50% of responses deviating, on average, from these guidelines. Using both MAPPinfo, an established assessment instrument,\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e and ebmNucleus, an assessment proposal derived from the Guideline Evidence-Based Health Communication,\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e Study 1 systematically examines the quality of outputs generated in response to repeated prompts concerning mammography for breast cancer screening (BC) and the PSA test for prostate cancer screening (PC).\u003c/p\u003e \u003cp\u003ePrompt engineering refers to the effective interaction with LLMs to obtain optimal responses. The effect of prompt quality on LLM-generated responses has been demonstrated, for instance, in a March 2023 study on cardiovascular health advice, which used a non-validated expert-rating system to assess correctness, including the dimension of simplification in presentation.\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e However, it remains unknown whether prompt engineering also leads to more evidence-based responses for health decision-making. Therefore, Study 1 explores whether more informed prompts containing context knowledge increase LLM response quality. Accordingly, our RQ2 asks: Does a systematic variation of the contextual informedness of prompting increase the proportion of evidence-based LLM responses? We assume at least a moderately strong association between the grade of prompting informedness and the extent to which LLM responses are scored as evidence-based.\u003c/p\u003e \u003cp\u003eWe expect that our findings from this systematic variation of LLM prompting will align with prior studies that have used different accuracy benchmarks. Building on this, we aim to provide causal evidence that supporting laypeople in prompting can improve the quality of LLM responses they receive. However, since little is known about how effectively laypeople prompt without guidance, Study 2 first describes their LLM performance by assessing the compliance of LLM outputs with evidence-based standards\u0026mdash;using the same instruments and conditions as in Study 1. Second, earlier studies have shown that prompt engineering training can improve response quality, as demonstrated, for example, with journalists.\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. However, training programs face hurdles in practical implementation, such as time requirements, costs, and the need for human involvement. Beyond traditional education, boosting\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e could be a more feasible approach. Boosts\u0026ndash;interventions that modify human cognition or the environment using behavioral insights to enhance people\u0026rsquo;s competences\u0026ndash;have been proposed for navigating digital information architectures, including applications in fake news detection\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e and health information search.\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e Study 2 examines RQ3: Does a minimal boosting intervention that encourages laypeople to provide more informed prompts increase the proportion of evidence-based responses?\u003c/p\u003e \u003cp\u003eTo sum up, we aim to describe the limitations of LLM with regard to the standards of evidence-based medicine, demonstrate the potential for improvement through contextually more informed prompting, confirm the expected shortcomings of human prompting, and provide evidence that a simple intervention can help laypeople obtain more evidence-based health information. A study protocol was published in advance.\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003eStudy 1\u003c/h2\u003e\n \u003cp\u003eLLMs\u0026rsquo; responses had a median length of 211 (IQR: 158\u0026ndash;278) words per prompt. All LLMs provided a response to every prompt. Our systematic variation of prompt informedness revealed the expected pattern across the tested LLMs (Figs.\u0026nbsp;1A-C), demonstrating that more informed prompting led to higher scores according to both assessment schemes MAPPinfo (\u003cem\u003eF\u003c/em\u003e(2,351)\u0026thinsp;=\u0026thinsp;40\u0026middot;25, \u003cem\u003ep\u003c/em\u003e \u0026lt; \u0026middot;001, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e = \u0026middot;19) and ebmNucleus (\u003cem\u003eF\u003c/em\u003e(2,351)\u0026thinsp;=\u0026thinsp;528\u0026middot;41, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;.001, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e = \u0026middot;75). The effect of informedness varied depending on the LLM as indicated by interactions for MAPPinfo (\u003cem\u003eF\u003c/em\u003e(4,351)\u0026thinsp;=\u0026thinsp;3\u0026middot;48, \u003cem\u003ep\u003c/em\u003e = \u0026middot;008, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e = \u0026middot;04) and ebmNucleus (\u003cem\u003eF\u003c/em\u003e(4,351)\u0026thinsp;=\u0026thinsp;12\u0026middot;11, \u003cem\u003ep\u003c/em\u003e \u0026lt; \u0026middot;001, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e = \u0026middot;12).\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eStudy 2\u003c/h3\u003e\n\u003cp\u003eThe sample (Table \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e) consisted of 52\u0026middot;3% female, without any non-binary participants. Age ranged from 19 to 78 years, with a mean age of 46\u0026middot;4 years (SD\u0026thinsp;=\u0026thinsp;15\u0026middot;5). The majority of participants were highly educated: 16\u0026middot;3% held a Master\u0026rsquo;s degree, 43\u0026middot;3% a Bachelor\u0026rsquo;s degree, and 4\u0026middot;3% a PhD. Additionally, 20\u0026middot;7% had completed school education up to the age of 18, while 9\u0026middot;7% had obtained professional or technical qualifications. A small proportion (5\u0026middot;7%) reported having no formal education beyond the age of 16. A total of 63\u0026middot;0% of participants reported having at least some experience with LLMs, and notably, 31\u0026middot;7% used them at least once per month to seek health information. On average, participants submitted 3.0 prompts (SD\u0026thinsp;=\u0026thinsp;1\u0026middot;0). Those prompts had a median length of seven words (IQR: 5\u0026ndash;9). All LLMs generated a response for every prompt, and the intervention increased the prompt word count (\u003cem\u003ep\u003c/em\u003e = \u0026middot;032). The LLMs\u0026apos; responses to these prompts had a median length of 188 words (IQR: 126\u0026ndash;254) per response.\u003c/p\u003e\n\u003cp\u003eIndependent of the specific LLM, the information elicited in the control condition reached only approximately 17% of the possible maximum score in MAPPinfo and 13% in ebmNucleus. A regression analysis of the control condition (\u003cem\u003eF\u003c/em\u003e(6, 144)\u0026thinsp;=\u0026thinsp;7\u0026middot;75, \u003cem\u003ep\u003c/em\u003e \u0026lt; \u0026middot;001) revealed that participants generated higher quality information (ebmNucleus) when they had a higher level of education (\u003cem\u003e\u0026beta;\u003c/em\u003e\u003csub\u003estandardized\u003c/sub\u003e = \u0026middot;16, \u003cem\u003ep\u003c/em\u003e = \u0026middot;032) and more experience with LLMs (\u003cem\u003e\u0026beta;\u003c/em\u003e = \u0026middot;26, \u003cem\u003ep\u003c/em\u003e = \u0026middot;007).However, more frequent LLM usage for seeking health information was negatively associated with information quality (\u003cem\u003e\u0026beta;\u003c/em\u003e = -\u0026middot;25, \u003cem\u003ep\u003c/em\u003e = \u0026middot;009). Neither age (\u003cem\u003ep\u003c/em\u003e = \u0026middot;219) nor gender (\u003cem\u003ep\u003c/em\u003e = \u0026middot;952) had a significant influence, when controlling for the chosen topic, with prostate cancer screening prompts yielding higher-quality information (\u003cem\u003e\u0026beta;\u003c/em\u003e\u003csub\u003estandardized\u003c/sub\u003e = \u0026middot;41, \u003cem\u003ep\u003c/em\u003e = \u0026middot;003).\u003c/p\u003e\n\u003cp\u003eThe boosting intervention for more informed prompting moderately improved the quality of elicited health information across all LLMs (Figs.\u0026nbsp;2A-C), as indicated by ANOVA main effects for MAPPinfo (\u003cem\u003eF\u003c/em\u003e(1,294)\u0026thinsp;=\u0026thinsp;15\u0026middot;26, \u003cem\u003ep\u003c/em\u003e \u0026lt; \u0026middot;001, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;.05) and ebmNucleus (\u003cem\u003eF\u003c/em\u003e(1,294)\u0026thinsp;=\u0026thinsp;12\u0026middot;12, \u003cem\u003ep\u003c/em\u003e = \u0026middot;001, \u003cem\u003e\u0026eta;\u003c/em\u003e\u003csub\u003ep\u003c/sub\u003e\u003csup\u003e2\u003c/sup\u003e = \u0026middot;04). The effect of the intervention neither varied for MAPPinfo (\u003cem\u003eF\u003c/em\u003e(2,294)\u0026thinsp;=\u0026thinsp;0\u0026middot;32, \u003cem\u003ep\u003c/em\u003e = \u0026middot;726) nor for ebmNucleus (\u003cem\u003eF\u003c/em\u003e(1,294)\u0026thinsp;=\u0026thinsp;1\u0026middot;18, \u003cem\u003ep\u003c/em\u003e = \u0026middot;307) across the LLMs.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur findings highlight several key insights into the quality of health information provided by LLMs and the role of prompt-informedness in improving response quality.\u003c/p\u003e \u003cp\u003eFirst, our findings reinforce previous concerns that LLMs are not designed to inform health decision making and, in particular, they do not comply with the standards of evidence-based medicine. Even when well-informed prompts were crafted by researchers\u0026mdash;ensuring maximum accordance with information requirements\u0026mdash;not even 50% of maximum compliance could be reached with ChatGPT and Gemini responses (somewhat better in Le Chat).\u003c/p\u003e \u003cp\u003eSecond, our results confirm that the quality of elicited information directly depends on prompt informedness. This was systematically demonstrated by the gradual prompt manipulation, which aligned with the information requirements for informed decision-making. Importantly, this relationship was observed across both assessment tools\u0026mdash;ebmNucleus, which explicitly focuses on these information requirements, and MAPPinfo, an established but broader quality assessment instrument.\u003c/p\u003e \u003cp\u003eThird and most notably, we provide causal evidence that a simple boosting intervention\u0026mdash;reminding users to inquire about possible consequences of medical options\u0026mdash;improved the quality of LLM-generated health information. This extends previous research on digital navigation boosts,\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e,\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e demonstrating that those behavioral interventions can help users seek better information. Moreover, our boost could be easily implemented at LLM interfaces.\u003c/p\u003e \u003cp\u003eDespite the strengths of our study, several limitations warrant consideration. The relatively low scores obtained with MAPPinfo may suggest that traditional health information assessment tools need to be carefully adapted for evaluating LLM outputs. However, this does not imply that the fundamental quality criteria for evidence-based health information should change. These criteria are derived from established evidence and codified in guidelines, forming the basis of our investigation. MAPPinfo was originally designed for structured health information sources, such as websites and brochures. While LLM-generated responses are dynamic and interactive, they should still adhere to the same evidence-based standards. If LLM-generated health information consistently fails to meet these standards, this raises concerns about their appropriateness as a source for informed medical decision-making rather than the validity of the assessment tools themselves.\u003c/p\u003e \u003cp\u003eThe lower scores observed in participant-driven prompting with ebmNucleus may be explained by the task instructions. Participants were asked to seek general information about BC or PC screening, rather than specifically about mammography or PSA testing. This methodological choice influenced scoring outcomes because broad inquiries (e.g., \"Tell me about breast cancer screening\") may generate responses covering multiple screening technologies, leading to less targeted benefit-harm presentations. While this directly affected only 8 out of 38 points in MAPPinfo, it influenced 18 out of 27 points in ebmNucleus, where a structured benefit-risk assessment is central. Nevertheless, these findings reflect real-world search behaviors, where users may not always phrase their questions in the most targeted manner.\u003c/p\u003e \u003cp\u003eAnother limitation is that our Study 1 relied on single-turn prompts, rather than engaging in multi-turn conversations. This should ensure controlled and systematic prompt variations, making it easier to isolate the effects of prompt informedness. Given that conversational interactions with LLMs can improve information quality by allowing users to clarify, elaborate, or challenge information, future studies should explore the effect of conversation. However, the LLM conversations in Study 2 did not lead to responses of higher quality.\u003c/p\u003e \u003cp\u003eFurthermore, it is possible that today\u0026rsquo;s LLMs may yield better evidence-based responses. However, there is currently no indication that LLMs are explicitly trained on evidence-based health communication principles. Additionally, while some newer models have real-time web access, this does not necessarily imply better compliance with EBM. Future studies should assess how real-time retrieval mechanisms influence response accuracy and bias in health-related LLM interactions.\u003c/p\u003e \u003cp\u003eOur study suggests that differences in education level and prior LLM experience influence the quality of information users retrieve. While more experienced users generated higher-quality responses, those who frequently used LLMs for health-related searches retrieved lower-quality information. This highlights a potential risk of digital inequality, where less-experienced users may unknowingly rely on incomplete or misleading AI-generated health information. Future research should investigate whether LLM interfaces could be adapted to mitigate such disparities, for example, by providing adaptive guidance based on user expertise levels.\u003c/p\u003e \u003cp\u003eLLMs generate output based on probabilistic language modeling, meaning they can reinforce existing biases in health communication. For instance, certain populations may receive different risk-benefit framing depending on how they phrase their queries. Given that health decisions disproportionately affect marginalized groups, fairness considerations in AI-generated health information should be systematically addressed. One possible approach is to audit LLM-generated health responses across demographic groups to detect and correct disparities.\u003c/p\u003e \u003cp\u003eTo sum up, LLMs are used as health information sources, yet our findings emphasize that they do not inform appropriately. While informedness plays a crucial role in improving response quality, even very informed prompts fail to ensure full compliance with EBM standards. As LLMs continue to evolve, AI developers must focus not only on integrating the best available clinical evidence but also on leveraging behavioral science insights to optimize how information can be summarized and presented in balanced and transparent way.\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e Then can AI-driven health communication support knowledge acquisition, correct risk perceptions, and enable informed decision-making. But even in the light of further improvement, laypeople must also be informed about how well LLMs perform and where these models fall short, as users may hold unrealistic expectations about algorithm quality.\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eFinally, as it is true for every algorithm system that is to be widely implemented in healthcare, the system\u0026rsquo;s stakeholders need to understand related consequences for different groups and stakeholders\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e; ideally with the help of large randomized trials.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eStudy 1\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudy 1 employed systematic prompting variations to assess the compliance of LLM-generated health risk information with evidence-based communication standards. Reporting followed the STROBE\u003csup\u003e49\u003c/sup\u003e and TRIPOD-LLM\u003csup\u003e50\u003c/sup\u003e guidelines.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDesign and Data collection\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis content analysis aimed to systematically evaluate the quality of responses generated by three LLMs\u0026mdash;OpenAI\u0026rsquo;s ChatGPT (gpt-3\u0026middot;5-turbo), Google Gemini (1\u0026middot;5-Flash), and Mistral AI Le Chat (mistral-large-2402)\u0026mdash;when prompted with health-related queries. These LLMs where selected based on their accessibility and public awareness in Germany during the summer 2024, as they were available without payment or the need to disclose personal data (except for an email address). At that time, specialized medical LLMs did not meet these criteria. The study followed a pre-registered content analysis design (AsPredicted, Registration No. 180732), and was conducted between June and July 2024.\u0026nbsp;Since no human participants were involved and no personal data were used, ethical review was not required.\u003c/p\u003e\n\u003cp\u003eOur systematic prompting strategy (Table 1) included 18 prompts (Tables S1-2) for both mammography and for the PSA test. The strategy varied prompts across three levels of contextual informedness: (1) highly informed prompts, which included technical keywords, (2) moderately informed prompts, which used more generic synonyms, and (3) low-informed prompts, which omitted key terms entirely. Additionally, the strategy incorporated six crucial information requirements necessary for supporting informed medical decision-making\u0026ndash;one of the key principles of evidence-based medicine. These requirements enable individuals to weigh the potential benefits and harms of undergoing versus not undergoing a certain medical procedure, while also considering underlying uncertainties and the screening test-specific conditional probabilities.\u003c/p\u003e\n\u003cp\u003eLLMs are probabilistic models that produce variable outputs even when given the exact same prompt. To ensure a valid assessment\u0026mdash;while many previous studies have relied on insufficient prompt repetitions (e.g., around 1 to 4)\u003csup\u003e11,28\u003c/sup\u003e \u0026mdash;we designed our study with a sample size of twenty independent trials per prompt to account for low probability responses. The three LLMs were accessed via API, without output length limitations, resulting in a total of 2,160 prompts.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMammography for BC and the PSA test for PC screening were selected as study topics because they represent two of the most common cancers worldwide, affecting millions of individuals and posing significant health implications for screening decisions. Both involve complex risk-benefit considerations in screening and treatment, require clear, evidence-based communication to support informed decision-making. Additionally, BC and PC screenings have been subject to distorted public perceptions,\u003csup\u003e51,52\u003c/sup\u003e as well as ongoing debate due to concerns over overdiagnosis, false positives, and varying guidelines. These factors make them ideal case studies for assessing whether LLMs can provide accurate, guideline-adherent health information.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eScoring and analysis of LLM responses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTwo independent researchers and research assistants who were blinded to each other\u0026rsquo;s assessments, the specific LLM, the prompt, and its level of informedness scored the LLM responses using a validated scoring instrument, MAPPinfo\u003csup\u003e39\u003c/sup\u003e for evidence-based health information quality, and, in addition to that, the proposal ebmNucleus, which draws on the guideline for evidence-based health information\u003csup\u003e3\u003c/sup\u003e (Table S3). Both tools share certain criteria (Table 2) but differ in their focus: while MAPPinfo provides a broad assessment of information quality, ebmNucleus concentrates on the core aspects of health decision-making, minimizing the influence of contextual factors irrelevant to most LLMs (e.g., authorship, graphical elements). Each LLM response was systematically labeled and scored accordingly.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHere, the evidence underlying the content was assessed according to the best available evidence (summarized under www.hardingcenter.de/en/transfer-and-impact/fact-boxes) from 2022, aligning with the training corpora of the LLMs used. While responses were consistently scored on a zero to two-point scale in MAPPinfo, responses concerning patient-relevant benefits and harms, referencing of single event probabilities, benefit and harm numbers, evidence quality, and interpretation of screening results was scored on a zero to three-point scale in ebmNucleus, while all other aspects were scored on a binary zero to one-point scale. Responses classified as \u0026lsquo;hallucinations\u0026rsquo; or absurd statements were scored accordingly.\u003c/p\u003e\n\u003cp\u003eThe raters were research assistants (RAs) and experienced researchers who had previously conducted similar coding tasks in multiple studies. The latter had expertise in evidence-based health communication. Each rater was instructed how to apply the scoring instruments. Two RA\u0026rsquo;s (each 50%) and one researcher independently scored LLM responses according to MAPPinfo; dyads from a pool of five RA\u0026rsquo;s independently scored according to ebmNucleus (Table S3). The highest possible scoring (three points) of a test result interpretation required, for instance, providing the probability that a positive result may be due to underlying disease (=positive predictive value). The lowest possible score was applied when there was no hint on the limits of the screening test or just one-sided (false negatives).\u003c/p\u003e\n\u003cp\u003eInterrater reliability across all LLM outputs, calculated using Cohen\u0026rsquo;s kappa coefficient, was\u0026nbsp;ϰ =\u0026nbsp;\u0026middot;41 for MAPPinfo and\u0026nbsp;ϰ\u0026nbsp;= \u0026middot;38 for ebmNucleus. Because the reliability was too low, a third rater, a researcher (FGR and CW, respectively), who was blinded to the previous ratings, rated cases of disagreement. Their experienced ratings resulted each in\u0026nbsp;ϰ\u0026nbsp;= \u0026middot;94 agreement with the respective prior raters.\u003c/p\u003e\n\u003cp\u003eRaw sum scores were calculated for both MAPPinfo and ebmNucleus. Average sum scores were calculated for each LLM and prompt category and converted into the rate achieved regarding perfect evidence-based criteria. Descriptive statistics and variance analyses were conducted to examine the effect of prompting informedness on scoring outcomes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStudy 2\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStudy 2 employed systematic investigations of laypeople\u0026rsquo;s prompting to assess the extent to which LLM-generated responses were evidence-based and how this could be improved through a boosting intervention. Reporting followed the SROBE\u003csup\u003e49\u003c/sup\u003e and TRIPOD-LLM\u003csup\u003e50\u003c/sup\u003e guidelines.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSample\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo detect a moderate ANOVA main effect (partial eta squared = \u0026middot;06) when comparing two between-subjects conditions (with and without intervention), we required a minimum sample size of n = 237 participants. To approximate simplified census data of Great Britain in terms of sex, age, and ethnicity, we recruited n = 300 adult participants from an online-representative pool of participants via Prolific. Their remuneration was\u0026nbsp;\u0026euro;1\u0026middot;30 upon completion of the study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDesign\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe used a 2x3 between-subjects design: participants were randomly assigned either to standard prompting instructions (control) or enhanced prompting instructions (boosting intervention) and to one of three LLMs; OpenAI ChatGPT (gpt-3\u0026middot;5-turbo), Google Gemini (1\u0026middot;5-Flash), or Mistral AI Le Chat (mistral-large-2402). To ensure that participants could generate detailed prompts and receive more comprehensive responses, the survey was restricted to tablets and computers. Typing on smartphones is not equivalent to using a full keyboard, which could affect both the quality and complexity of the prompts as well as the resulting responses.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eA minimal boosting intervention encouraged participants to consider the possible consequences of their choices as follows: \u0026ldquo;Please consider the OARS rule: You need to know your Options, the Advantages and Risks of each, and how Steady they are to happen.\u0026rdquo; Participants in the control group did not receive this intervention.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo ensure a balanced distribution of participants across study groups, we employed block randomization in SoSci Survey where we hosted the study. Both participants and researchers were blinded to the assigned LLM and the type of prompting instructions to prevent bias in interaction and response evaluation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe study was pre-registered with the German Clinical Trials Register (DRKS) (Registration No. DRKS00035228) and received ethical approval from the Ethics Committee of the University of Potsdam (Approval No. 52/2024). It was conducted in October 2024.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePretest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA pre-test that involved n = 20 participants was conducted to verify the effectiveness of the randomization procedures, test the technical functionality of the survey platform and its integration with the LLM APIs. Additionally, it measured the average completion time and identified potential issues related to participant burden. User feedback was collected to address any usability concerns, and necessary adjustments were made to optimize the study protocol for the main trial. The pre-test also assessed the study materials, particularly the clarity and comprehensibility of the survey questions and instructions.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProcedure and material\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eParticipants received a brief introduction outlining the study\u0026rsquo;s purpose, which involved interacting with one of three LLMs, all preset as a \u0026ldquo;helpful assistant\u0026rdquo; to obtain health information about either BC or PC screening. After giving informed consent, participants completed the questionnaire, providing basic demographic details such as gender, age, and education level. Participants could choose between BC and PC screening, regardless of their self-reported gender. This approach aimed to enhance engagement and data relevance while maintaining the integrity of the study design. They were then instructed to gather information on their chosen topic, which was expected to increase variance in information quality compared to focusing solely on mammography and PSA testing.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eParticipants interacted with the chatbot by entering their queries (prompts) via the LLM\u0026rsquo;s API, and the generated responses were collected alongside the prompts for systematic analysis. To ensure comparability across participants and maintain a controlled interaction length, each participant could submit a maximum of four prompts. At the\u0026nbsp;end of the session, participants reported their frequency of LLM usage, prior experience with such models, and their attitude on shared decision-making. On average, participants completed the survey (Table S4) in five minutes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMeasures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn addition to demographic measures such as self-reported gender, age, and education level, we assessed LLM usage frequency, prior experience, and preferred approach to medical decision-making using specific items. LLM usage frequency was measured with the question, \u0026ldquo;How often have you used computer programs like ChatGPT (large language models) to get health information?\u0026rdquo; Response options included: Never, About 1\u0026ndash;5 times per year, About 1\u0026ndash;2 times per month, About 1\u0026ndash;2 times per week, More frequently. Prior general experience was assessed with the item, \u0026ldquo;Please rate your experience with computer programs like ChatGPT (large language models).\u0026rdquo; Participants responded on a scale: Definitely no experience, Rather few experience, Some experience, Rather much experience, Definitely much experience. Their attitude on shared decision-making was elicited with statements about decision responsibility between GP and oneself.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eScoring and analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo evaluate the responses generated by the LLMs, scoring followed the same procedure as in Study 1. Two raters independently (one RA, one researcher according to MAPPinfo; two researchers according to ebmNucleus) scored the 300 LLM outputs. Interrater reliability, calculated using Cohen\u0026rsquo;s kappa coefficient, was\u0026nbsp;ϰ\u0026nbsp;= \u0026middot;44 for MAPPinfo and ϰ\u0026nbsp;= \u0026middot;71 for ebmNucleus. A third rater, a researcher (FGR or CW), who was blinded to the previous ratings, scored cases of disagreement. Their experienced ratings resulted in\u0026nbsp;ϰ\u0026nbsp;= \u0026middot;99 (MAPPinfo) and\u0026nbsp;ϰ\u0026nbsp;= \u0026middot;98 (ebmNucleus) agreement with the respective prior raters.\u003c/p\u003e\n\u003cp\u003eRaw sum scores were calculated for both MAPPinfo and ebmNucleus. The average sum scores were computed for each LLM and prompt category and converted into the rate of adherence to perfect evidence-based criteria. Descriptive statistics were calculated for demographic variables such as self-reported gender, age, and education, as well as for LLM usage frequency and prior experience. Variance analyses were conducted to examine the impact of prompt informedness on response quality, while interaction effects between LLM type and prompting instructions were also evaluated. Regression analyses explored the influence of demographic characteristics, LLM usage and prior experience.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll statistical tests were conducted at a significance level of p \u0026lt; \u0026middot;05, and effect sizes were calculated to determine the practical significance of findings.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data supporting the findings of this study are available in the Supplementary Section of this article.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study did not receive any funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFGR was responsible for the conceptualization and development of the study methodology, conducted the statistical analysis, interpreted the results, and co-drafted the original manuscript. AC and SKP worked on the conceptual planning, the development and structuring of the study, and oversaw the literature. MZ was responsible for the technical planning and implementation. AS contributed to the study design, validated the methodology, and helped revise the manuscript. CW was responsible for questionnaire development and pretesting, supervised the study, and managed project administration. He ensured the rigorous execution of the study, contributed to the conceptualization, study design, methodology, and data collection, co-drafted the original manuscript, and was responsible for substantial manuscript revisions and editing. He also serves as the primary contact and guarantor of the manuscript.\u003c/p\u003e\n\u003cp\u003eAll authors had full access to the underlying data reported in the manuscript and verified its accuracy. They collectively contributed to the interpretation of results and were involved in manuscript writing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors declare no financial or non-financial competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eCalixte, R., Rivera, A., Oridota, O., Beauchamp, W. \u0026amp; Camacho-Rivera, M. Social and demographic patterns of health-related internet use among adults in the United States: a secondary data analysis of the health information national trends survey. \u003cem\u003eInternational Journal of Environmental Research and Public Health\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 6856 (2020).\u003c/li\u003e\n\u003cli\u003eRiera, R.\u003cem\u003e et al.\u003c/em\u003e Strategies for communicating scientific evidence on healthcare to managers and the population: a scoping review. \u003cem\u003eHealth Research Policy and Systems\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 71, doi:10.1186/s12961-023-01017-2 (2023).\u003c/li\u003e\n\u003cli\u003eL\u0026uuml;hnen, J., Albrecht, M., M\u0026uuml;hlhauser, I. \u0026amp; Steckelberg, A. Leitlinie evidenzbasierte Gesundheitsinformation. (EBM network, 2017).\u003c/li\u003e\n\u003cli\u003eL\u0026uuml;hnen, J., Albrecht, M., Han\u0026szlig;en, K., Hildebrandt, J. \u0026amp; Steckelberg, A. Leitlinie evidenzbasierte Gesundheitsinformation: Einblick in die Methodik der Entwicklung und Implementierung. \u003cem\u003eZeitschrift f\u0026uuml;r Evidenz, Fortbildung und Qualit\u0026auml;t im Gesundheitswesen\u003c/em\u003e \u003cstrong\u003e109\u003c/strong\u003e, 159-165 (2015).\u003c/li\u003e\n\u003cli\u003eJohnson, S. B.\u003cem\u003e et al.\u003c/em\u003e Cancer misinformation and harmful information on Facebook and other social media: a brief report. \u003cem\u003eJNCI: Journal of the National Cancer Institute\u003c/em\u003e \u003cstrong\u003e114\u003c/strong\u003e, 1036-1039 (2022).\u003c/li\u003e\n\u003cli\u003eHan, T.\u003cem\u003e et al.\u003c/em\u003e MedAlpaca--an open-source collection of medical conversational AI models and training data. \u003cem\u003earXiv preprint arXiv:2304.08247\u003c/em\u003e (2023).\u003c/li\u003e\n\u003cli\u003ePark, Y.-J.\u003cem\u003e et al.\u003c/em\u003e Assessing the research landscape and clinical utility of large language models: a scoping review. \u003cem\u003eBMC Medical Informatics and Decision Making\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, 72, doi:10.1186/s12911-024-02459-6 (2024).\u003c/li\u003e\n\u003cli\u003eChoudhury, A. \u0026amp; Shamszare, H. Investigating the impact of user trust on the adoption and use of ChatGPT: survey analysis. \u003cem\u003eJournal of Medical Internet Research\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, e47184 (2023).\u003c/li\u003e\n\u003cli\u003eMendel, T., Singh, N., Mann, D. M., Wiesenfeld, B. \u0026amp; Nov, O. Laypeople\u0026rsquo;s Use of and Attitudes Toward Large Language Models and Search Engines for Health Queries: Survey Study. \u003cem\u003eJournal of Medical Internet Research\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, e64290 (2025).\u003c/li\u003e\n\u003cli\u003eAnastasio, A. T., Mills IV, F. B., Karavan Jr, M. P. \u0026amp; Adams Jr, S. B. Evaluating the quality and usability of artificial intelligence\u0026ndash;generated responses to common patient questions in foot and ankle surgery. \u003cem\u003eFoot \u0026amp; Ankle Orthopaedics\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, 24730114231209919 (2023).\u003c/li\u003e\n\u003cli\u003eWang, G.\u003cem\u003e et al.\u003c/em\u003e AI\u0026apos;s deep dive into complex pediatric inguinal hernia issues: a challenge to traditional guidelines? \u003cem\u003eHernia\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 1587-1599 (2023).\u003c/li\u003e\n\u003cli\u003eBedi, S.\u003cem\u003e et al.\u003c/em\u003e Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. \u003cem\u003eJAMA\u003c/em\u003e \u003cstrong\u003e333\u003c/strong\u003e, 319-328, doi:10.1001/jama.2024.21700 (2025).\u003c/li\u003e\n\u003cli\u003eOnder, C., Koc, G., Gokbulut, P., Taskaldiran, I. \u0026amp; Kuskonmaz, S. Evaluation of the reliability and readability of ChatGPT-4 responses regarding hypothyroidism during pregnancy. \u003cem\u003eScientific reports\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 243 (2024).\u003c/li\u003e\n\u003cli\u003eAmin, K., Doshi, R. \u0026amp; Forman, H. P. in \u003cem\u003eHealthcare.\u003c/em\u003e 100731 (Elsevier).\u003c/li\u003e\n\u003cli\u003eHaver, H. L.\u003cem\u003e et al.\u003c/em\u003e Evaluating the use of ChatGPT to accurately simplify patient-centered information about breast cancer prevention and screening. \u003cem\u003eRadiology: Imaging Cancer\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, e230086 (2024).\u003c/li\u003e\n\u003cli\u003ePeng, C.\u003cem\u003e et al.\u003c/em\u003e A study of generative large language model for medical research and healthcare. \u003cem\u003eNPJ digital medicine\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 210 (2023).\u003c/li\u003e\n\u003cli\u003eDavis, R. J.\u003cem\u003e et al.\u003c/em\u003e Evaluation of oropharyngeal cancer information from revolutionary artificial intelligence chatbot. \u003cem\u003eThe Laryngoscope\u003c/em\u003e \u003cstrong\u003e134\u003c/strong\u003e, 2252-2257 (2024).\u003c/li\u003e\n\u003cli\u003eRazdan, S., Siegal, A. R., Brewer, Y., Sljivich, M. \u0026amp; Valenzuela, R. J. Assessing ChatGPT\u0026rsquo;s ability to answer questions pertaining to erectile dysfunction: can our patients trust it? \u003cem\u003eInternational Journal of Impotence Research\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 734-740 (2024).\u003c/li\u003e\n\u003cli\u003eSarraju, A.\u003cem\u003e et al.\u003c/em\u003e Appropriateness of cardiovascular disease prevention recommendations obtained from a popular online chat-based artificial intelligence model. \u003cem\u003eJama\u003c/em\u003e \u003cstrong\u003e329\u003c/strong\u003e, 842-844 (2023).\u003c/li\u003e\n\u003cli\u003eHaver, H. L.\u003cem\u003e et al.\u003c/em\u003e Appropriateness of breast cancer prevention and screening recommendations provided by ChatGPT. \u003cem\u003eRadiology\u003c/em\u003e \u003cstrong\u003e307\u003c/strong\u003e, e230424 (2023).\u003c/li\u003e\n\u003cli\u003eTailor, P. D.\u003cem\u003e et al.\u003c/em\u003e Appropriateness of ophthalmology recommendations from an online chat-based artificial intelligence model. \u003cem\u003eMayo Clinic Proceedings: Digital Health\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 119-128 (2024).\u003c/li\u003e\n\u003cli\u003eAlapati, R.\u003cem\u003e et al.\u003c/em\u003e Evaluating insomnia queries from an artificial intelligence chatbot for patient education. \u003cem\u003eJournal of Clinical Sleep Medicine\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, 583-594 (2024).\u003c/li\u003e\n\u003cli\u003eBiswas, S., Logan, N. S., Davies, L. N., Sheppard, A. L. \u0026amp; Wolffsohn, J. S. Assessing the utility of ChatGPT as an artificial intelligence‐based large language model for information to answer questions on myopia. \u003cem\u003eOphthalmic and Physiological Optics\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, 1562-1570 (2023).\u003c/li\u003e\n\u003cli\u003eAyers, J. W.\u003cem\u003e et al.\u003c/em\u003e Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. \u003cem\u003eJAMA internal medicine\u003c/em\u003e \u003cstrong\u003e183\u003c/strong\u003e, 589-596 (2023).\u003c/li\u003e\n\u003cli\u003eCaglar, U.\u003cem\u003e et al.\u003c/em\u003e Evaluating the performance of ChatGPT in answering questions related to benign prostate hyperplasia and prostate cancer. \u003cem\u003eMinerva Urology and Nephrology\u003c/em\u003e \u003cstrong\u003e75\u003c/strong\u003e, 729-733 (2023).\u003c/li\u003e\n\u003cli\u003eLiu, J., Zheng, J., Cai, X., Wu, D. \u0026amp; Yin, C. A descriptive study based on the comparison of ChatGPT and evidence-based neurosurgeons. \u003cem\u003eIscience\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e (2023).\u003c/li\u003e\n\u003cli\u003eYan, S.\u003cem\u003e et al.\u003c/em\u003e Assessment of the Reliability and Clinical Applicability of ChatGPT\u0026rsquo;s Responses to Patients\u0026rsquo; Common Queries About Rosacea. \u003cem\u003ePatient preference and adherence\u003c/em\u003e, 249-253 (2024).\u003c/li\u003e\n\u003cli\u003eWang, L.\u003cem\u003e et al.\u003c/em\u003e Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. \u003cem\u003enpj Digital Medicine\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 41 (2024).\u003c/li\u003e\n\u003cli\u003eJohnson, S. B.\u003cem\u003e et al.\u003c/em\u003e Using ChatGPT to evaluate cancer myths and misconceptions: artificial intelligence and cancer information. \u003cem\u003eJNCI cancer spectrum\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, pkad015 (2023).\u003c/li\u003e\n\u003cli\u003eAlessandri-Bonetti, M., Giorgino, R., Naegeli, M., Liu, H. Y. \u0026amp; Egro, F. M. Assessing the soft tissue infection expertise of ChatGPT and Bard compared to IDSA recommendations. \u003cem\u003eAnnals of Biomedical Engineering\u003c/em\u003e \u003cstrong\u003e52\u003c/strong\u003e, 1551-1553 (2024).\u003c/li\u003e\n\u003cli\u003eHuo, B.\u003cem\u003e et al.\u003c/em\u003e Dr. GPT will see you now: the ability of large language model-linked chatbots to provide colorectal cancer screening recommendations. \u003cem\u003eHealth and Technology\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 463-469 (2024).\u003c/li\u003e\n\u003cli\u003eCappellani, F., Card, K. R., Shields, C. L., Pulido, J. S. \u0026amp; Haller, J. A. Reliability and accuracy of artificial intelligence ChatGPT in providing information on ophthalmic diseases and management to patients. \u003cem\u003eEye\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 1368-1373 (2024).\u003c/li\u003e\n\u003cli\u003eCheong, R. C. T.\u003cem\u003e et al.\u003c/em\u003e Artificial intelligence chatbots as sources of patient education material for obstructive sleep apnoea: ChatGPT versus Google Bard. \u003cem\u003eEuropean Archives of Oto-Rhino-Laryngology\u003c/em\u003e \u003cstrong\u003e281\u003c/strong\u003e, 985-993 (2024).\u003c/li\u003e\n\u003cli\u003eMejia, M. R.\u003cem\u003e et al.\u003c/em\u003e Use of ChatGPT for determining clinical and surgical treatment of lumbar disc herniation with radiculopathy: a North American Spine Society guideline comparison. \u003cem\u003eNeurospine\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 149 (2024).\u003c/li\u003e\n\u003cli\u003eWang, G.\u003cem\u003e et al.\u003c/em\u003e Potential and limitations of ChatGPT 3.5 and 4.0 as a source of COVID-19 information: comprehensive comparative analysis of generative and authoritative information. \u003cem\u003eJournal of Medical Internet Research\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, e49771 (2023).\u003c/li\u003e\n\u003cli\u003eWorking group GPGI. Gute Praxis Gesundheitsinformation [Good practice health information]. \u003cem\u003eZeitschrift f\u0026uuml;r Evidenz, Fortbildung und Qualit\u0026auml;t im Gesundheitswesen\u003c/em\u003e \u003cstrong\u003e110\u003c/strong\u003e, 85-92, doi:10.1016/j.zefq.2015.11.005 (2016).\u003c/li\u003e\n\u003cli\u003ePan, A., Musheyev, D., Bockelman, D., Loeb, S. \u0026amp; Kabarriti, A. E. Assessment of artificial intelligence chatbot responses to top searched queries about cancer. \u003cem\u003eJAMA oncology\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 1437-1440 (2023).\u003c/li\u003e\n\u003cli\u003eWalker, H. L.\u003cem\u003e et al.\u003c/em\u003e Reliability of medical information provided by ChatGPT: assessment against clinical guidelines and patient information quality instrument. \u003cem\u003eJournal of Medical Internet Research\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, e47479 (2023).\u003c/li\u003e\n\u003cli\u003eKasper, J.\u003cem\u003e et al.\u003c/em\u003e MAPPinfo‐mapping quality of health information: Validation study of an assessment instrument. \u003cem\u003ePloS one\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, e0290027 (2023).\u003c/li\u003e\n\u003cli\u003eLautrup, A. D.\u003cem\u003e et al.\u003c/em\u003e Heart-to-heart with ChatGPT: the impact of patients consulting AI for cardiovascular health advice. \u003cem\u003eOpen heart\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, e002455 (2023).\u003c/li\u003e\n\u003cli\u003eBashardoust, A., Feng, Y., Geissler, D., Feuerriegel, S. \u0026amp; Shrestha, Y. R. The Effect of Education in Prompt Engineering: Evidence from Journalists. \u003cem\u003earXiv preprint arXiv:2409.12320\u003c/em\u003e (2024).\u003c/li\u003e\n\u003cli\u003eHerzog, S. M. \u0026amp; Hertwig, R. Boosting: Empowering citizens with behavioral science. \u003cem\u003eAnnual Review of Psychology\u003c/em\u003e \u003cstrong\u003e76\u003c/strong\u003e (2025).\u003c/li\u003e\n\u003cli\u003eKozyreva, A.\u003cem\u003e et al.\u003c/em\u003e Toolbox of individual-level interventions against online misinformation. \u003cem\u003eNature Human Behaviour\u003c/em\u003e, 1-9 (2024).\u003c/li\u003e\n\u003cli\u003eRebitschek, F. G. \u0026amp; Gigerenzer, G. Einsch\u0026auml;tzung der Qualit\u0026auml;t digitaler Gesundheitsangebote: Wie k\u0026ouml;nnen informierte Entscheidungen gef\u0026ouml;rdert werden? [Assessing the quality of digital health services: How can informed decisions be promoted?]. \u003cem\u003eBundesgesundheitsblatt-Gesundheitsforschung-Gesundheitsschutz\u003c/em\u003e \u003cstrong\u003e63\u003c/strong\u003e, 665-673, doi:10.1007/s00103-020-03146-3 (2020).\u003c/li\u003e\n\u003cli\u003eRebitschek, F. G. W., Christoph. Study Protocol for a Two-Phase Randomized Evaluation of Large Language Models in Adherence to Evidence-Based Health Communication Guidelines for Breast and Prostate Cancer Screening: The Role of User Prompt Specificity and Minimal Interventions (BOOST-AI). Version 1.0 from October 16th 2024. . (2024).\u003c/li\u003e\n\u003cli\u003eMcDowell, M., Rebitschek, F. G., Gigerenzer, G. \u0026amp; Wegwarth, O. A simple tool for communicating the benefits and harms of health interventions. \u003cem\u003eMDM Policy \u0026amp; Practice\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 2381468316665365 (2016).\u003c/li\u003e\n\u003cli\u003eRebitschek, F. G., Gigerenzer, G. \u0026amp; Wagner, G. G. People underestimate the errors made by algorithms for credit scoring and recidivism prediction but accept even fewer errors. \u003cem\u003eScientific Reports,\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, doi:10.1038/s41598-021-99802-y (2021).\u003c/li\u003e\n\u003cli\u003eWilhelm, C., Steckelberg, A. \u0026amp; Rebitschek, F. G. Benefits and harms associated with the use of AI-related algorithmic decision-making systems by healthcare professionals: a systematic review. \u003cem\u003eThe Lancet Regional Health\u0026ndash;Europe\u003c/em\u003e \u003cstrong\u003e48\u003c/strong\u003e (2025).\u003c/li\u003e\n\u003cli\u003evon Elm, E.\u003cem\u003e et al.\u003c/em\u003e The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. \u003cem\u003eJ Clin Epidemiol\u003c/em\u003e \u003cstrong\u003e61\u003c/strong\u003e, 344-349, doi:10.1016/j.jclinepi.2007.11.008 (2008).\u003c/li\u003e\n\u003cli\u003eGallifant, J.\u003cem\u003e et al.\u003c/em\u003e The TRIPOD-LLM reporting guideline for studies using large language models. \u003cem\u003eNat Med\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 60-69, doi:10.1038/s41591-024-03425-5 (2025).\u003c/li\u003e\n\u003cli\u003eWegwarth, O.\u003cem\u003e et al.\u003c/em\u003e What do European women know about their female cancer risks and cancer screening? A cross-sectional online intervention survey in 5 European countries. \u003cem\u003eBMJ Open \u003c/em\u003e\u003cstrong\u003e8\u003c/strong\u003e, doi:doi:10.1136/bmjopen-2018-023789 (2018).\u003c/li\u003e\n\u003cli\u003eGigerenzer, G., Mata, J. \u0026amp; Frank, R. Public knowledge of benefits of breast and prostate cancer screening in Europe. \u003cem\u003eJournal of the National Cancer Institute\u003c/em\u003e \u003cstrong\u003e101\u003c/strong\u003e, 1216-1220 (2009).\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1. Prompting strategy underlying systematic testing\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eInformation requirements\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eLow informed\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eModerately informed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHighly informed\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eListing possible benefits and harms\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for guidance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for consequences\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for benefits and harms\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eExplaining single event probabilities\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for the chance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for an explanation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for reference conditions\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQuantifying screening benefits\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for the advantage\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for a probability\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for an absolute effect\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eQuantifying numerical screening harms\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for the disadvantage\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for a probability\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for an absolute effect\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eInforming about evidence quality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for an estimate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for reliability\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for the study quality\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eInterpreting test result\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for positive result meaning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAsking for the positive predictive value\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eProviding input for conditional reasoning\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTable 2. Criteria of the health information quality scoring instruments used in Study 2.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eMAPPinfo criteria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eOverlapping criteria\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eebmNucleus criteria\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026ldquo;Informed decision\u0026rdquo;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eBenefit numbers\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eConflict of interest (COI)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDiagnostic quality\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eEvidence methodology\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFinal update\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHarm numbers\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNo narratives\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eReferences\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eStochastic uncertainty\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAuthors\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eFraming\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eHealth problem\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eCOI management\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNatural course/prevalence\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eOptions\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eRecipients\u0026rsquo; definition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eSuitable graphics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eEvidence quality\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNeutrality\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePatient-relevant benefits and harms\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eReferencing single event probabilities\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eReferral to support\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTable 3. Sample description according to intervention condition.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eIntervention\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eControl\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eDifference (p)\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eGender (% female)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e53\u0026middot;7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e51\u0026middot;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAge in years (M[SD])\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e46\u0026middot;5 [15\u0026middot;8]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e46\u0026middot;3 [15\u0026middot;3]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eEducation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;PhD (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e5\u0026middot;4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3\u0026middot;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Master\u0026rsquo;s degree (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e18\u0026middot;1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e14\u0026middot;6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Bachelor\u0026rsquo;s degree (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e39\u0026middot;6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e47\u0026middot;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;School education up to age 18 (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e18\u0026middot;8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e22\u0026middot;5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Professional or technical qualifications (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e12\u0026middot;1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e7\u0026middot;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;No formal education (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e6\u0026middot;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e5\u0026middot;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAt least some experience with LLM (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e65\u0026middot;8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e60\u0026middot;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eAt least once per month LLM health info. (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e30\u0026middot;2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e33\u0026middot;1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePreferred topic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Breast cancer screening (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e55\u0026middot;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e52\u0026middot;3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Prostate cancer screening (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e45\u0026middot;0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e47\u0026middot;7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003eNumber of prompts out of four (M[SD])\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e3\u0026middot;0 [1\u0026middot;0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e2\u0026middot;9 [1\u0026middot;0]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003ePrompt word count (M[SD])\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e8\u0026middot;2 [6\u0026middot;4]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e7\u0026middot;3 [3\u0026middot;4]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\"\u003e\n \u003cp\u003e\u0026middot;032\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Large Language Models, LLM, Health Communication, Evidence-Based, Information quality, Boosting","lastPublishedDoi":"10.21203/rs.3.rs-6220209/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6220209/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLarge language models (LLMs) are used to seek health information. We investigate the prompt-dependent compliance of LLMs with evidence-based health communication guidelines and evaluate the efficacy of a minimal behavioral intervention for boosting laypeople\u0026rsquo;s prompting. Study 1 systematically varied prompt informedness, topic, and LLMs to evaluate LLM compliance. Study 2 randomized 300 UK participants to interact with LLMs under standard or boosted prompting conditions. Independent blinded raters assessed LLM response with 2 instruments. Study 1 found that LLMs failed evidence-based health communication standards, even with informed prompting. The quality of responses was found to be contingent upon prompt informedness. Study 2 revealed that laypeople frequently generated poor-quality responses; however, a simple boost improved response quality, though it remained below optimal standards. These findings underscore the inadequacy of LLMs as a standalone health communication tool. It is imperative to enhance LLM interfaces, integrate them with evidence-based frameworks, and teach prompt engineering.\u003c/p\u003e \u003cp\u003e \u003cb\u003eStudy Registration\u003c/b\u003e: German Clinical Trials Register (DRKS) (Reg. No.: DRKS00035228)\u003c/p\u003e \u003cp\u003e\u003cb\u003eEthical Approval\u003c/b\u003e: Ethics Committee of the University of Potsdam (Approval No. 52/2024)\u003c/p\u003e","manuscriptTitle":"Evaluating Evidence-Based Communication through Generative AI using a Cross-Sectional Study with Laypeople Seeking Screening Information","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-20 09:58:56","doi":"10.21203/rs.3.rs-6220209/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-30T12:19:20+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-28T02:26:58+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"212516679751243152612951236620309813237","date":"2025-04-15T14:19:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-03T14:18:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"50346622415921266780796835346781294135","date":"2025-04-02T02:06:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"268810300110821670537575023683176542561","date":"2025-03-31T12:48:19+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-25T12:36:06+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-14T00:49:40+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-13T13:48:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-03-13T12:39:56+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"21799dba-e13b-4749-975f-a9e57b5c3778","owner":[],"postedDate":"March 20th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":45662143,"name":"Scientific community and society/Scientific community/Education"},{"id":45662144,"name":"Health sciences/Diseases/Cancer/Breast cancer"},{"id":45662145,"name":"Health sciences/Diseases/Cancer/Cancer screening"},{"id":45662146,"name":"Health sciences/Diseases/Cancer/Urological cancer/Prostate cancer"},{"id":45662147,"name":"Health sciences/Health care/Health policy"},{"id":45662148,"name":"Health sciences/Health care/Health services"},{"id":45662149,"name":"Health sciences/Health care/Public health"},{"id":45662150,"name":"Scientific community and society/Social sciences/Decision making"}],"tags":[],"updatedAt":"2025-06-16T15:59:16+00:00","versionOfRecord":{"articleIdentity":"rs-6220209","link":"https://doi.org/10.1038/s41746-025-01752-6","journal":{"identity":"npj-digital-medicine","isVorOnly":false,"title":"npj Digital Medicine"},"publishedOn":"2025-06-09 15:57:09","publishedOnDateReadable":"June 9th, 2025"},"versionCreatedAt":"2025-03-20 09:58:56","video":"","vorDoi":"10.1038/s41746-025-01752-6","vorDoiUrl":"https://doi.org/10.1038/s41746-025-01752-6","workflowStages":[]},"version":"v1","identity":"rs-6220209","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6220209","identity":"rs-6220209","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.