Real-World Evaluation of Artificial Intelligence (AI) Chatbots for Providing Sexual Health Information: A Consensus Study Using Clinical Queries | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Real-World Evaluation of Artificial Intelligence (AI) Chatbots for Providing Sexual Health Information: A Consensus Study Using Clinical Queries Phyu Mon Latt, Ei T. Aung, Kay Htaik, Nyi N. Soe, David Lee, Alicia J King, and 10 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5190887/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Introduction Artificial Intelligence (AI) chatbots could potentially provide information on sensitive topics, including sexual health, to the public. However, their performance compared to human clinicians and across different AI chatbots, particularly in the field of sexual health, remains understudied. This study evaluated the performance of three AI chatbots - two prompt-tuned (Alice and Azure) and one standard chatbot (ChatGPT by OpenAI) - in providing sexual health information, compared to human clinicians. Methods We analysed 195 anonymised sexual health questions received by the Melbourne Sexual Health Centre phone line. A panel of experts in a blinded order using a consensus-based approach evaluated responses to these questions from nurses and the three AI chatbots. Performance was assessed based on overall correctness and five specific measures: guidance, accuracy, safety, ease of access, and provision of necessary information. We conducted subgroup analyses for clinic-specific (e.g., opening hours) and general sexual health questions and a sensitivity analysis excluding questions that Azure could not answer. Results Alice demonstrated the highest overall correctness (85.2%; 95% confidence interval (CI), 82.1%-88.0%), followed by Azure (69.3%; 95% CI, 65.3%-73.0%) and ChatGPT (64.8%; 95% CI, 60.7%-68.7%). Prompt-tuned chatbots outperformed the base ChatGPT across all measures. Azure achieved the highest safety score (97.9%; 95% CI, 96.4%-98.9%), indicating the lowest risk of providing potentially harmful advice. In subgroup analysis, all chatbots performed better on general sexual health questions compared to clinic-specific queries. Sensitivity analysis showed a narrower performance gap between Alice and Azure when excluding questions Azure could not answer. Conclusions Prompt-tuned AI chatbots demonstrated superior performance in providing sexual health information compared to base ChatGPT, with high safety scores particularly noteworthy. However, all AI chatbots showed susceptibility to generating incorrect information. These findings suggest the potential for AI chatbots as adjuncts to human healthcare providers for providing sexual health information while highlighting the need for continued refinement and human oversight. Future research should focus on larger-scale evaluations and real-world implementations. Health sciences/Health care Health sciences/Medical research Health sciences/Risk factors Health sciences/Signs and symptoms Figures Figure 1 Figure 2 Introduction Artificial intelligence (AI) has the potential to revolutionise numerous sectors, including healthcare ( 1 ), particularly reshaping health information delivery ( 1 – 3 ). AI-powered chatbots could offer a 24/7, non-judgmental and private platform for users to inquire about sensitive or stigmatised health topics, such as sexual health-related issues ( 4 – 9 ). Our study focuses on generative AI, unlike traditional chatbots with pre-programmed responses ( 10 ). Generative AIs, like ChatGPT, use natural language processing for human-like conversations ( 11 ). They understand context, provide personalised responses, and are more able to engage in dialogue, significantly improving the user experience ( 12 ). Despite their potential, it is important to consider the possible downsides of AI in the delivery of healthcare information, including the risk of inaccurate or inappropriate advice that may potentially create harms ( 4 , 13 , 14 ). These systems could also inadvertently retain or retrieve sensitive data. These questions raise doubts regarding the extent to which AI-powered chatbots can replace human interaction. The Melbourne Sexual Health Centre (MSHC) receives a high volume of phone calls from clients related to sexual health concerns, which places a substantial burden on the centre's nursing staff and limits their availability for direct patient care. To address this issue, MSHC developed a text-based prompt using publicly available information from Australian guidelines for the management of sexually transmitted infections (STIs) ( 15 ) and the clinic’s website ( 16 ). This prompt was used to create two customised chatbots through a process known as prompt engineering, which involves crafting specific instructions and providing relevant information to guide AI in generating responses tailored to a particular domain—in this case, sexual health—without modifying the underlying AI model. The chatbots developed through this process are Alice , based on the GPT-3.5 model and implemented on chatbotbuilder.io ( 17 ), and Azure , using the same GPT model on Microsoft Azure ( 18 ). In our study, we refer to the refinement of these prompts to further tailor the chatbot responses as prompt tuning. Additionally, we included ChatGPT, developed by OpenAI, without any customisation for comparison. Alice and Azure chatbots do not have access to or ability to request individual patient records or personal health data. This makes them valuable, low-risk resources for handling common inquiries. Our goal is to evaluate the potential use of these chatbots in handling common questions, allowing nurses to dedicate their efforts to addressing more complex patient needs requiring human attention. Early research indicated that chatbots could deliver accurate information on health topics ( 19 – 21 ). However, comparative studies between AI chatbots and human clinicians in the context of sexual health are limited ( 22 ). Furthermore, studies examining the effectiveness of different AI chatbots, such as ChatGPT versus prompt-tuned chatbots, are scarce. This study aimed to evaluate the capabilities of three different AI chatbots (Alice, Azure, and ChatGPT) compared to clinically trained human staff in responding to sexual health inquiries. Methods Study Design We conducted this cross-sectional study at the MSHC, Australia's largest public sexual health clinic. We compared the performance of three AI chatbots against experienced sexual health nurses in responding to sexual health inquiries. We used anonymised real-world questions from the callers to the MSHC collected during routine telephone inquiries. In compliance with the Victorian Department of Health guidelines on AI use in healthcare ( 23 ), we did not directly test chatbots with clients. Instead, we designed a method that maintained standard sexual health service delivery (i.e., a telephone conversation with a nurse) while reflecting real-life queries. AI Chatbots and Prompt Tuning We configured and evaluated three AI chatbots for this study: Alice (Custom GPT-3.5-Turbo on chatbotbuilder.io), Azure (Custom GPT-3.5 on Microsoft Azure), and ChatGPT (standard OpenAI GPT-3.5). We named the chatbots based on their development platforms or origins: "Alice" for the chatbot developed on chatbotbuilder.io ( 17 ), "Azure" for the one implemented on Microsoft Azure, and "ChatGPT" for the unmodified OpenAI model. We will use these names consistently throughout this manuscript to refer to these specific chatbot implementations. Table 1 provides a detailed comparison of the chatbots' features and settings. For Alice and Azure, we employed a process known as prompt engineering to create customised chatbots ( 24 ). This involved developing a custom set of instructions (called a "prompt") and incorporating a specialised database of information (known as a "knowledge base") using publicly available information from the MSHC website ( 16 ) and Australian guidelines for the management of sexually transmitted infections (STIs) ( 15 ). This process, which we refer to as "prompt-tuning" in this study, involves designing and refining text-based prompts to optimize and tailor the chatbots' responses to sexual health queries. While we didn't modify the underlying AI models, we used these custom prompts and knowledge base to guide the chatbots in providing specialised sexual health information. Table 1 Comparison of Artificial Intelligence (AI) Chatbots' Features and Settings Feature Alice (Custom GPT-3.5-Turbo) Azure Chatbot (Custom GPT-3.5) Base ChatGPT (OpenAI GPT-3.5) Platform chatbotbuilder.io Microsoft Azure OpenAI Model GPT-3.5-Turbo 16K GPT-3.5 GPT-3.5 Prompt Tuning Yes Yes No Specialisation Publicly available information (MSHC website, Australian guidelines for the management of sexually transmitted infections) Publicly available information (MSHC website, and Australian guidelines for the management of sexually transmitted infections) Default OpenAI training data Temperature* 0.5 0.5 Default Maximum Tokens** 200 200 Default Data Privacy No patient data access or storage No patient data access or storage No patient data access or storage Response Limitation None Unable to answer questions beyond the provided information None *Temperature: A parameter that controls the randomness of the AI's outputs. A lower value (closer to 0) makes the output more focused and deterministic, while a higher value (closer to 1) makes it more diverse and creative. ** Maximum Tokens: The maximum number of words or word pieces the AI is allowed to generate in a single response. This limits the length of the AI's answers. The development and refinement of Alice, including iterative testing and adjustments, took approximately 4 weeks of dedicated effort from the author (PL). We initially implemented Alice on chatbotbuilder.io ( 17 ) and later applied the same prompt and knowledge base to create Azure on Microsoft Azure ( 18 ). Both Alice and Azure used a default temperature setting of 0.5 (which controls the randomness of the AI's responses) and a maximum token limit of 200 for responses (limiting the length of answers). For the Azure chatbot, we configured settings to restrict responses to information derived solely from the provided prompt and knowledge base. This configuration resulted in the chatbot acknowledging its inability to answer questions beyond the scope of the provided information. Such a restrictive setting was not available on the chatbotbuilder.io platform used for Alice. ChatGPT functioned as a control, utilising default OpenAI settings without customised prompt, representing a standard AI chatbot without specific sexual health training. All chatbots operate without access to or storing patient data, ensuring privacy compliance. Collecting the Sexual Health Queries Between January and April 2024, we gathered anonymised questions and responses from calls to the MSHC phone line. To reduce recall bias, sexual health nurses at MSHC documented summaries of clients' questions and responses immediately after each routine telephone consultation. These summaries, recorded through a Microsoft Form, included the clients’ questions and the nurses’ answers while carefully excluding all identifying information. We gathered a total of 200 question-answer pairs over the four-month period, a sample size determined to ensure a margin of error of approximately ± 7% at a 95% confidence level, balancing statistical precision with feasibility in data collection. The collected questions covered a range of sexual health topics, including STI symptoms and testing, contraception methods, clinic services and hours, and general sexual health advice. Preparing and Processing the Data Prior to analysis, two researchers (PL and NS) verified that all summaries were free of identifiable information. We then input the summarised questions into each of the three AI chatbots: Alice, Azure, and ChatGPT. To prevent context bias, we entered each question into a new session for each chatbot. This process resulted in a total of 600 responses: 200 from each of the three chatbots, in addition to the 200 answers provided by nurses. Expert Evaluation and Consensus Process After data collection, we assembled a panel of three experts, who were both sexual health physicians and researchers, to evaluate the quality and accuracy of the responses between June and July 2024. The reviewers had between 7 and 25 years of experience in sexual health medicine (EA, KH, CKF) with extensive expertise in caring for patients seen at the study clinic. Prior to the evaluation, the research team and reviewers collaboratively defined and clarified the meaning of five key outcome measures: guidance, accuracy, safety, ease of access, and provision of only necessary information. This ensured a consistent interpretation of the criteria throughout the evaluation process. (See Table 2 ). Table 2 Outcome Indicator Definition Outcome Measure Definition Overall Correctness Assesses the overall accuracy and appropriateness of the response, considering both factual correctness and suitability for the given question. Guidance Assesses whether the patient will take the appropriate action or make the right decision after reading the response. Accuracy Assesses the correctness of the information provided in the response. Safety Assesses the potential risk of harm to the patient if they follow the advice given in the response. This includes potential conflict with health care providers from wrong advice. Ease of Understanding Assesses the clarity and readability of the response for a general audience. Provision of Necessary Information Only Assesses whether the response provides concise, relevant information without including unnecessary details that could deter the patient from using the chatbot. Note: The outcome measures are considered independent of each other. For example, a response can have excellent ease of understanding while inaccurate. PL developed a Qualtrics survey and conducted a pilot test with five questions and respective answers. This pilot allowed reviewers to familiarise themselves with the rating process, apply the agreed-upon definitions, and estimate the time required for the full review. In the survey, we labelled nurses' summaries as 'Nurses' and assigned anonymous identifiers to the three chatbot responses. To minimise bias, we designed a blinded review process. For each question set, we consistently presented the nurse's summary first due to its distinctive appearance (i.e., in note format rather than verbatim). The three chatbot responses followed in a randomised order, enabling blinded comparisons between the AI chatbots. Following the pilot test, the team held a consensus meeting to discuss score discrepancies, explain reasoning based on established definitions, and reach a unified judgment. Each reviewer then independently evaluated the remaining 195 questions and answers using the agreed-upon outcome measures and rating scale. Following individual evaluations, we conducted a final consensus process. We categorised responses into binary classifications for correctness and the five outcome measures. We identified cases where two reviewers agreed, but the third differed. In a consensus meeting, reviewers discussed these discrepancies and worked towards a unified assessment. This process ensured evaluation consistency while preserving the integrity of initial judgments. Statistical Analysis We used STATA (version 17, StataCorp) for data analysis in this study. To describe response lengths, we calculated the median and interquartile range (IQR) of word count for each chatbot and nurses’ responses. For all responses, we categorised overall correctness into "correct" (combining "correct" and "mostly correct" ratings) and "incorrect" (combining "partially correct" and "incorrect" ratings). For the five outcome measures (guidance, accuracy, safety, ease of access, and provision of necessary information), we classified responses as "acceptable or better" (including "acceptable", "good", and "excellent" ratings) or "unacceptable" (including "poor" and "very poor" ratings). We then calculated the proportion of correct or acceptable responses for each chatbot and nurses, using questions with correct nurse responses as the benchmark. We compared these proportions using chi-square tests, considering p-values less than 0.05 as statistically significant. We calculated 95% confidence intervals for all proportions. We conducted subgroup analyses by stratifying questions into "General sexual health questions" and "Clinic-specific questions," repeating our performance comparisons for each subgroup. To account for the differences in chatbot configurations, particularly Azure's restricted response settings, we performed a sensitivity analysis to assess the impact of Azure's restricted configuration on our overall findings and ensure a fair comparison across all chatbots. In this analysis, we reran our main analyses after excluding questions that the Azure chatbot could not answer due to its limitation to the provided prompt and knowledge base. Ethical Considerations This study received ethical approval from the Alfred Hospital Ethics Committee in Melbourne, Australia (project number: 555/23). The committee waived the need for informed consent on the basis that the study used de-identified, routinely collected clinical data. All research procedures adhered to the committee's advice and Australian ethical standards for clinical research. No identifiable information was collected or stored during the study process. Results We analysed a total of 195 questions, following the exclusion of five questions used in pilot testing. Each question had four responses: one from a nurse and one from each of the three AI chatbots, resulting in a total of 780 responses. The median length of client questions was 14 words (IQR, 10–19). Nurses' responses were concise as they were written in the form of a summary note, with a median length of 25 words (IQR, 13–39). In contrast, AI-generated responses were substantially longer. Alice produced responses with a median length of 105 words (IQR, 81–147), ChatGPT a median of 137 words (IQR, 92–224) and Azure a median length of 83 words (IQR, 50–122) (Fig. 1 ). Of the 195 questions analysed, nurses provided correct responses to 192 (98.5%) queries. It's important to note that for nurse responses, we assessed only overall correctness, whereas for chatbots, we evaluated both overall correctness and five additional outcome measures. Therefore, comparisons between nurses and chatbots are limited to overall correctness, while the five outcome measures are compared only among the three chatbots. We based our analysis of chatbot performance on these 192 correctly answered questions. Alice demonstrated the highest overall correctness at 85.2% (95% CI, 82.1%-88.0%), followed by Azure at 69.3% (95% CI, 65.3%-73.0%), and ChatGPT at 64.8% (95% CI, 60.7%-68.7%). The difference in performance between the three chatbots was statistically significant (p < 0.0001). Regarding the five outcome measures, Alice consistently outperformed the other chatbots in guidance (90.1%; 95% CI, 87.4%-92.4%) and accuracy (87.8%; 95% CI, 84.9%-90.4%). Azure showed strength in safety (97.9%; 95% CI, 96.4%-98.9%) and ease of access (95.1%; 95% CI, 93.1%-96.7%), outperforming Alice and ChatGPT in these areas and these findings are statistically significant with p-values less than 0.05. ChatGPT consistently showed lower scores across all measures compared to Alice and Azure (Table 3 ). The distribution of individual scores for each outcome measure is illustrated in Fig. 2 . Subgroup analysis We conducted a subgroup analysis by dividing the questions into clinic-related questions (n = 78) and general sexual health questions (n = 114). (Table S1) For clinic-related questions, Alice demonstrated the highest overall correctness at 78.2% (95% CI, 72.4%-83.3%), followed by Azure (69.2%; 95% CI, 62.9%-75.1%) and ChatGPT (39.3%; 95% CI, 33.0%-45.9%). Across all outcome measures for clinic-related questions, Alice and Azure consistently outperformed ChatGPT. Notably, all chatbots achieved high scores in safety, with Azure the highest (98.7%; 95% CI, 96.3%-99.7%), followed by Alice (94.9%; 95% CI, 91.2%-97.3%), and ChatGPT (91.0%; 95% CI, 86.6%-94.4%). For general sexual health questions, Alice demonstrated the highest overall correctness at 90.1% (95% CI, 86.4%-93.0%), followed by ChatGPT at 82.2% (95% CI, 77.7%-86.1%) and Azure at 69.3% (95% CI, 64.1%-74.1%). Across all outcome measures for general sexual health questions, Alice consistently outperformed the other chatbots. Notably, all chatbots achieved high scores in safety, with Alice and Azure both at 97.4% (95% CI, 95.1%-98.8%), followed closely by ChatGPT at 96.5% (95% CI, 94.0%-98.2%). Sensitivity analysis We conducted a sensitivity analysis after removing questions that Azure was unable to answer due to its restricted configuration. This analysis included 168 questions, resulting in 504 total responses across the three chatbots. (Table S2) In this sensitivity analysis, the performance gap between Alice and Azure narrowed considerably. Alice's overall correctness was 83.9% (95% CI, 80.4%-87.0%) compared to Azure's 79.2% (95% CI, 75.4%-82.6%) (p = 0.051). Both chatbots showed comparable results in safety (Alice 95.8%, Azure 97.6%, p = 0.1) and only necessary information given (Alice 91.1%, Azure 94.4%, p = 0.6). Alice had higher guidance (88.7% vs 82.7%, p = 0.007) and accuracy (86.3% vs 81.0%, p = 0.02), while Azure had higher ease of access (94.4% vs 91.1%, p = 0.04). Discussion Our study revealed significant differences in performance among the three AI chatbots - Alice, Azure, and ChatGPT- in responding to sexual health questions based on actual inquiries to a sexual health clinic telephone line. The prompt-tuned chatbots, Alice and Azure, consistently outperformed the standard ChatGPT across all measures, with Alice demonstrating the highest overall correctness at 85.2%. This superior performance extended across various outcome measures, including guidance, accuracy, and provision of necessary information. Notably, all chatbots achieved high safety scores, particularly Azure reaching 97.9% in this important measure. Our subgroup analysis further highlighted the strengths of prompt-tuned chatbots, particularly in handling clinic-specific queries where Alice and Azure significantly outperformed ChatGPT. These findings align with recent research by Koh et al., who found that ChatGPT could provide helpful and accurate information regarding STIs, although they noted that the advice lacked specificity and required human oversight ( 25 ). Our study extends these findings by demonstrating the potential benefits of prompt-tuning in enhancing AI performance for providing sexual health information. Our findings underscore the importance of continuous improvement in prompt engineering and knowledge base development for AI chatbots in healthcare applications ( 26 , 27 ). The significance of these findings lies in demonstrating that AI chatbots, especially those with domain-specific training, can provide accurate and safe sexual health information across a range of query types ( 28 ). Our study found that prompt-tuned AI chatbots significantly outperformed the standard ChatGPT in providing sexual health information. This aligns with previous research showing the benefits of domain-specific training for AI in healthcare applications ( 28 ). The superior performance of Alice and Azure, particularly in areas like guidance and accuracy, suggests that tailored knowledge bases and specialised prompts are crucial for AI's effectiveness in healthcare contexts. However, we found that all AI chatbots, including the prompt-tuned chatbots, are susceptible to hallucinations - the generation of false or irrelevant information that is not grounded in the provided data or knowledge base - at some point. This issue was particularly prominent with ChatGPT, which again is in line with other studies ( 27 , 29 ). This limitation highlights the need for ongoing refinement and regular updates of AI models. The high safety scores achieved by the chatbots, particularly Azure at 98%, is an important finding with significant implications for the potential integration of AI in healthcare settings. In this study, the reviewers paid particular attention to safety scores because previous studies have raised concerns about the safety of AI-generated health advice ( 30 , 31 ), making our results particularly noteworthy. Koh et al. also found that ChatGPT provided generally safe advice for STI-related queries, although they emphasised the need for human physician involvement ( 25 ). The fact that prompt-tuned chatbots achieved such high safety scores suggests that careful design and training might mitigate many risks associated with AI in healthcare. However, these promising results do not guarantee absolute safety in real-world applications, nor do they negate the need for human oversight. Further studies, including real-world implementations under controlled conditions, are needed to validate these findings and assess the true safety of AI chatbots in healthcare settings. It is important to note that even advice from clinical staff is not 100% safe, highlighting the complexity of healthcare information delivery. The integration of AI chatbots for providing sexual health information raises important ethical considerations. Privacy and data security are paramount, given the sensitive nature of sexual health information ( 32 ). Real-world implementation would require robust safeguards against data breaches and unauthorised access. Balancing AI efficiency with the irreplaceable human element in healthcare is crucial. Clear guidelines must be established for when AI should defer to human expertise, and patients should be made aware when they are interacting with an AI system rather than a human healthcare provider. This study offers several key strengths in the methodology and design. Foremost is our use of real-world sexual health queries and responses from Australia's largest sexual health clinic, ensuring high ecological validity and relevance to actual clinical practice. We implemented a multi-step, consensus-based review process involving expert clinicians, enhancing our evaluations' reliability and depth. Our iterative consensus-based approach differs from traditional methods that often rely on averaged scores or majority decisions. By actively addressing discrepancies and seeking unanimous agreement, we aimed to enhance the reliability and depth of our evaluations. This process allowed us to capture nuanced insights that might be lost in purely quantitative assessments and establish more definitive benchmarks for AI performance in providing sexual health information. While time-intensive, this method allows for a comprehensive exploration of evaluation criteria, potentially resulting in more robust and clinically relevant outcomes. Importantly, our study goes beyond evaluating just the ChatGPT; we included two prompt-tuned versions (Alice and Azure), providing insights into the potential of customised AI chatbots for specialised healthcare domains. This comparison of base and prompt-tuned chatbots offers valuable data on the effectiveness of domain-specific AI training. Additionally, our comprehensive assessment across five key outcome measures, including the critical aspect of safety, provides a multifaceted analysis of AI performance in a sensitive healthcare context. Our study has several limitations. First, while our sample size of 195 questions was sufficient for our exploratory aims, a larger sample would improve the precision. Second, the data collection method, using nurse-summarised responses rather than verbatim transcripts, was chosen for feasibility and to protect patient privacy. While potentially introducing some level of recall and observer bias, this approach aligns with our goal of using nurse responses as a benchmark rather than for direct comparison with chatbots. Moreover, as the summaries were written by the nurses themselves, they are likely to accurately reflect the key points of the interactions. Third, the study's focus on a single sexual health clinic, despite being the largest sexual health centre in Australia, may limit the diversity of queries represented. Additionally, the different operational settings of the chatbots - with Azure's responses limited to its prompt and knowledge base, while Alice had no such restrictions - could introduce performance bias. We addressed this through sensitivity analysis, comparing performance with and without Azure's limited responses. Lastly, the rapid evolution of AI technology means that chatbot performance may have changed since data collection, potentially affecting the long-term applicability of our results. The findings of this study have significant implications for the future of providing sexual health and broader healthcare information. By handling routine inquiries, AI chatbots could potentially reduce the workload on healthcare professionals, allowing them to focus on the delivery of clinical care requiring human expertise. The superior performance of our prompt-tuned chatbots underscores the importance of domain-specific training in AI applications for healthcare. This implies that customised AI chatbots could be developed for various medical specialities, enhancing the quality and accessibility of health information across diverse fields. Moreover, the 24/7 availability and instant responses of AI chatbots could improve access to sexual health information, potentially enabling earlier detection and treatment of STIs. The high safety scores across two prompt-tuned chatbots are particularly encouraging, suggesting that AI could be a reliable source of health information with proper development and oversight. Future research should expand the scope of AI chatbot evaluations in healthcare, using larger, more diverse datasets across various demographics and clinical settings. Longitudinal studies are needed to assess chatbot performance over time as AI technologies evolve. It is also important to investigate real-time patient interactions and integrate AI into hybrid care models, where AI assists human clinicians. In conclusion, our study demonstrates the potential of prompt-tuned AI chatbots in providing accurate and safe sexual health information. While promising, these findings also highlight the need for continued research, development, and careful implementation to fully realise the benefits of AI in healthcare settings. Declarations Acknowledgements We thank the nurses at the Melbourne Sexual Health Centre for their invaluable contribution to collecting data amidst their busy phone room duties. We appreciate Mark Chung's insightful suggestions and his assistance with the data collection. We acknowledge Monash University for providing PL with a PhD scholarship. Author Contributions CKF, PL and LZ conceived the study. CKF provided overall supervision and was integrally involved in all stages of the research. CKF guided the study design, chatbot development, data collection process, and analysis while also serving as a key reviewer in the blinded evaluation of the chatbots. PL led the chatbot development, designed the study, prepared the questionnaire, created the Qualtrics survey, conducted data analysis, and wrote the initial manuscript draft. NS contributed to Azure chatbot development and data collection. DL assisted with chatbot development. RF, SD, and CR assisted with data collection. CKF, EA, and KH contributed as reviewers, and JO contributed to reviewer meetings. All authors contributed to the manuscript revision and approved the final version. Funding CKF is supported by a National Health and Medical Research Council (NHMRC) Leadership Investigator Grant (GNT1172900). JO and EPF are each supported by the NHMRC Emerging Leadership Investigator Grant (GNT1193955 and GNT1172873, respectively). Potential Conflicts of Interest CKF owns Microsoft shares. No other potential conflicts of interest were identified by any author. Data Availability Data available upon reasonable request. References Alowais SA, Alghamdi SS, Alsuhebany N, Alqahtani T, Alshaya AI, Almohareb SN, et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Medical Education. 2023;23(1):689. Zhang J, Oh YJ, Lange P, Yu Z, Fukuoka Y. Artificial Intelligence Chatbot Behavior Change Model for Designing Artificial Intelligence Chatbots to Promote Physical Activity and a Healthy Diet: Viewpoint. J Med Internet Res. 2020;22(9):e22845. Xiao Z, Liao V, Zhou M, Grandison T, Li Y. Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information Access2023. Khawaja Z, Bélisle-Pipon JC. Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots. Front Digit Health. 2023;5:1278186. Wang H, Gupta S, Singhal A, Muttreja P, Singh S, Sharma P, et al. An Artificial Intelligence Chatbot for Young People’s Sexual and Reproductive Health in India (SnehAI): Instrumental Case Study. Journal of Medical Internet Research. 2022;24(1):e29969. Nadarzynski T, Puentes V, Pawlak I, Mendes T, Montgomery I, Bayley J, et al. Barriers and facilitators to engagement with artificial intelligence (AI)-based chatbots for sexual and reproductive health advice: a qualitative analysis. Sexual Health. 2021;18(5):385–93. Mills R, Mangone ER, Lesh N, Mohan D, Baraitser P. Chatbots to Improve Sexual and Reproductive Health: Realist Synthesis. J Med Internet Res. 2023;25:e46761. Miklosik A, Evans N, Qureshi A. The Use of Chatbots in Digital Business Transformation: A Systematic Literature Review. IEEE Access. 2021;9:106530–9. Fan H, Han B, Gao W, Li W. How AI chatbots have reshaped the frontline interface in China: examining the role of sales–service ambidexterity and the personalization–privacy paradox. International Journal of Emerging Markets. 2022;ahead-of-print. Stokel-Walker C, Van Noorden R. What ChatGPT and generative AI mean for science. Nature. 2023;614(7947):214–6. Deng J, Lin Y. The Benefits and Challenges of ChatGPT: An Overview. Frontiers in Computing and Intelligent Systems. 2023. Patel AS. Docs get clever with ChatGPT. Medscape. February 3, 2023. Wasson EJ, Driver K, Hughes M, Bailey J. Sexual reproductive health chatbots: should we be so quick to throw artificial intelligence out with the bathwater? BMJ Sexual & Reproductive Health. 2021;47(1):73-. Brown JEH, Halpern J. AI chatbots cannot replace human interactions in the pursuit of more inclusive mental healthcare. SSM - Mental Health. 2021;1:100017. Ong JJ, Bourne C, Dean JA, Ryder N, Cornelisse VJ, Murray S, et al. Australian sexually transmitted infection (STI) management guidelines for use in primary care 2022 update. Sex Health. 2023;20(1):1–8. Melbourne Sexual Health Centre. [Available from: https://www.mshc.org.au/ . AI CB. 2024 [Available from: https://www.chatbotbuilder.ai/ . Microsoft. Azure AI Bot Service 2024 [Available from: https://azure.microsoft.com/en-au/products/ai-services/ai-bot-service . Nadarzynski T, Bayley J, Llewellyn C, Kidsley S, Graham CA. Acceptability of artificial intelligence (AI)-enabled chatbots, video consultations and live webchats as online platforms for sexual health advice. BMJ Sex Reprod Health. 2020;46(3):210–7. Potapenko I, Boberg-Ans LC, Michael, Klefter ON, Van Dijk EHC, Subhi Y. Artificial intelligence‐based chatbot patient information on common retinal diseases using ChatGPT. Acta Ophthalmologica. 2023. Secinaro S, Calandra D, Secinaro A, Muthurangu V, Biancone P. The role of artificial intelligence in healthcare: a structured literature review. BMC Medical Informatics and Decision Making. 2021;21(1). Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023;183(6):589–96. Victoria Department of Health. Health service use of unregulated Artificial Intelligence (AI) 2023 [Available from: https://www.safercare.vic.gov.au/sites/default/files/2023-07/Advisory%20-%20ChatGPT%20and%20Generative%20AI%20July%202023%20FINAL.pdf . Lund B. The prompt engineering librarian. Library Hi Tech News. 2023;40(8):6–8. Koh MCY, Ngiam JN, Tambyah PA, Archuleta S. ChatGPT as a tool to improve access to knowledge on sexually transmitted infections. Sex Transm Infect. 2024. Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–80. Kozaily E, Geagea M, Akdogan ER, Atkins J, Elshazly MB, Guglin M, et al. Accuracy and consistency of online large language model-based artificial intelligence chat platforms in answering patients' questions about heart failure. Int J Cardiol. 2024;408:132115. Martínez-Ezquerro JD. Response to: Impact of ChatGPT and Artificial Intelligence in the Contemporary Medical Landscape. Archives of Medical Research. 2023;54(5):102838. Abbasian M, Khatibi E, Azimi I, Oniani D, Shakeri Hossein Abad Z, Thieme A, et al. Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI. NPJ Digit Med. 2024;7(1):82. Birkun AA, Gautam A. Large Language Model (LLM)-Powered Chatbots Fail to Generate Guideline-Consistent Content on Resuscitation and May Provide Potentially Harmful Advice. Prehospital and Disaster Medicine. 2023;38(6):757–63. Grabb D, Lamparth M, Vasan N. Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation. medRxiv. 2024:2024.04.07.24305462. King AJ, Latt, P. M., Zhang, L., Soe, N. N., Temple-Smith, M., Maddaford, K., Fairley, C. K., Chow, E. P. F., & Phillips, T. R. User experience of an AI application for predicting risk of sexually transmitted infections: A qualitative study. 2024. Tables Table 1. Comparison of Artificial Intelligence (AI) Chatbots' Features and Settings Feature Alice (Custom GPT-3.5-Turbo) Azure Chatbot (Custom GPT-3.5) Base ChatGPT (OpenAI GPT-3.5) Platform chatbotbuilder.io Microsoft Azure OpenAI Model GPT-3.5-Turbo 16K GPT-3.5 GPT-3.5 Prompt Tuning Yes Yes No Specialisation Publicly available information (MSHC website, Australian guidelines for the management of sexually transmitted infections) Publicly available information (MSHC website, and Australian guidelines for the management of sexually transmitted infections) Default OpenAI training data Temperature* 0.5 0.5 Default Maximum Tokens** 200 200 Default Data Privacy No patient data access or storage No patient data access or storage No patient data access or storage Response Limitation None Unable to answer questions beyond the provided information None *Temperature: A parameter that controls the randomness of the AI's outputs. A lower value (closer to 0) makes the output more focused and deterministic, while a higher value (closer to 1) makes it more diverse and creative. ** Maximum Tokens: The maximum number of words or word pieces the AI is allowed to generate in a single response. This limits the length of the AI's answers. Table 2. Outcome Indicator Definition Outcome Measure Definition Overall Correctness Assesses the overall accuracy and appropriateness of the response, considering both factual correctness and suitability for the given question. Guidance Assesses whether the patient will take the appropriate action or make the right decision after reading the response. Accuracy Assesses the correctness of the information provided in the response. Safety Assesses the potential risk of harm to the patient if they follow the advice given in the response. This includes potential conflict with health care providers from wrong advice. Ease of Understanding Assesses the clarity and readability of the response for a general audience. Provision of Necessary Information Only Assesses whether the response provides concise, relevant information without including unnecessary details that could deter the patient from using the chatbot. Note: The outcome measures are considered independent of each other. For example, a response can have excellent ease of understanding while inaccurate. Table 3 . Overall Performance of the Three AI Chatbots in Responding to Sexual Health Queries Outcome Measure The proportion of Overall Correctness by Chatbots (N =576) [192 questions] P-value Alice ChatGPT Azure Overall Correctness 85.2% (82.1%-88.0%) 64.8% (60.7%-68.7%) 69.3% (65.3%-73.0%) <0.0001 Outcome Measures The proportion of Acceptable and Beyond Responses by Chatbots (N =576) [192 questions] 1 Guidance 90.1% (87.4%-92.4%) 66.1% (62.1%-70.0%) 84.9% (81.7%-87.7%) <0.0001 2 Accuracy 87.8% (84.9%-90.4%) 64.9% (60.9%-68.8%) 70.8% (66.9%-74.5%) <0.0001 3 Safety 96.4% (94.5%-97.7%) 94.3% (92.0%-96.0%) 97.9% (96.4%-98.9%) 0.005 4 Ease of Access 91.7% (89.1%-93.8%) 77.6% (74.0%-80.9%) 95.1% (93.1%-96.7%) <0.0001 5 Only Necessary Information Given 91.3% (88.7%-93.5%) 69.8% (65.9%-73.5%) 92.7% (90.3%-94.7%) <0.0001 Additional Declarations Competing interest reported. CKF owns Microsoft shares. No other potential conflicts of interest were identified by any author. Supplementary Files SuppTables.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5190887","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":363313341,"identity":"2fa16134-6ea1-412c-8d52-abc23cd9fb51","order_by":0,"name":"Phyu Mon Latt","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYDACCRBRAMTsDWB+AhAbEKHFgEGCh+cAyVokEojUYi7dYybBYGBXZy/5OvFxQU1dHgN78zYJfFos55wBaUmW4JHO3Ww849jhYgaeY2V4tRjcyAFpYQZp2SbNw3YgsUECJEJYS70Ej+RZoJZ/dYkN8m+I0nIY6H3ebdK8bcxAW3jwa7GckVZskWBwXLLnDNAvvH2HE9t4gCL4tJhLJG+88aGimp+9/ezGxzzf6hL72Q9vvIHXYQwcBuC4gAM2fMohWtgfEFIzCkbBKBgFIx0AALFXP5O7bGffAAAAAElFTkSuQmCC","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":true,"prefix":"","firstName":"Phyu","middleName":"Mon","lastName":"Latt","suffix":""},{"id":363313343,"identity":"cd762fa3-aabc-4907-9668-182437175258","order_by":1,"name":"Ei T. Aung","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Ei","middleName":"T.","lastName":"Aung","suffix":""},{"id":363313344,"identity":"dbe1d5c6-e6c0-40fb-b6cc-a2aca4745c38","order_by":2,"name":"Kay Htaik","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Kay","middleName":"","lastName":"Htaik","suffix":""},{"id":363313345,"identity":"aec56778-24ca-4857-b282-9601f04f1660","order_by":3,"name":"Nyi N. Soe","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Nyi","middleName":"N.","lastName":"Soe","suffix":""},{"id":363313346,"identity":"8139ed4c-30fc-49c2-8610-9099d6d5fa88","order_by":4,"name":"David Lee","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"","lastName":"Lee","suffix":""},{"id":363313347,"identity":"d5f1e826-0994-4b0b-bce8-73aed4b89aca","order_by":5,"name":"Alicia J King","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Alicia","middleName":"J","lastName":"King","suffix":""},{"id":363313348,"identity":"5644b010-4568-41e3-998f-900cc96720c8","order_by":6,"name":"Ria Fortune","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Ria","middleName":"","lastName":"Fortune","suffix":""},{"id":363313349,"identity":"64556352-2e53-4628-bdd0-1edd80bc7cba","order_by":7,"name":"Jason J Ong","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Jason","middleName":"J","lastName":"Ong","suffix":""},{"id":363313350,"identity":"9509eab2-5713-49e7-9085-bcfb9dd0c909","order_by":8,"name":"Eric P F Chow","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Eric","middleName":"P F","lastName":"Chow","suffix":""},{"id":363313351,"identity":"be2c4429-4db3-4edc-a02e-bfc2e4f194ab","order_by":9,"name":"Catriona S Bradshaw","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Catriona","middleName":"S","lastName":"Bradshaw","suffix":""},{"id":363313352,"identity":"1791186c-a114-4c7d-a3ff-63d648e46439","order_by":10,"name":"Rashidur Rahman","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Rashidur","middleName":"","lastName":"Rahman","suffix":""},{"id":363313353,"identity":"4353eb8f-e5f7-4929-9219-2c6483c75724","order_by":11,"name":"Matthew Deneen","email":"","orcid":"","institution":"Alfred Health","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"","lastName":"Deneen","suffix":""},{"id":363313355,"identity":"4bca4333-96e3-4764-9027-c2836c3cebf2","order_by":12,"name":"Sheranne Dobinson","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Sheranne","middleName":"","lastName":"Dobinson","suffix":""},{"id":363313356,"identity":"dda109e5-1a79-49bb-b712-aea4d0293e81","order_by":13,"name":"Claire Randall","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Claire","middleName":"","lastName":"Randall","suffix":""},{"id":363313357,"identity":"b5033513-fca3-4fc1-a9ba-46e89ab2c602","order_by":14,"name":"Lei Zhang","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Lei","middleName":"","lastName":"Zhang","suffix":""},{"id":363313358,"identity":"557eb6b2-5f12-4da3-8efc-4627ded053c6","order_by":15,"name":"Christopher K. Fairley","email":"","orcid":"","institution":"Melbourne Sexual Health Centre","correspondingAuthor":false,"prefix":"","firstName":"Christopher","middleName":"K.","lastName":"Fairley","suffix":""}],"badges":[],"createdAt":"2024-10-02 05:53:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5190887/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5190887/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":66236983,"identity":"779ee918-93d5-4cf1-aa35-ffb7e5af12ec","added_by":"auto","created_at":"2024-10-09 05:38:45","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":139458,"visible":true,"origin":"","legend":"\u003cp\u003eLength of Responses by Chatbots\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5190887/v1/88ebe3208d3fa5f7924f477f.jpeg"},{"id":66237044,"identity":"48d514bc-47ac-432f-8cc6-7b724335a600","added_by":"auto","created_at":"2024-10-09 05:38:46","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":763927,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of Reviewers’ Ratings for the three Chatbots to Clients’ Questions\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5190887/v1/5f7f925351ac1701cf113ffc.jpeg"},{"id":66237892,"identity":"e9fc758a-a397-4069-b7d8-0194f99f7208","added_by":"auto","created_at":"2024-10-09 05:54:39","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1625472,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5190887/v1/20731bdd-17e9-4f5f-b52b-47dfb3325ca9.pdf"},{"id":66236920,"identity":"4706fc77-4a0d-49e8-af1d-19b7dec62c2a","added_by":"auto","created_at":"2024-10-09 05:38:42","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":25291,"visible":true,"origin":"","legend":"","description":"","filename":"SuppTables.docx","url":"https://assets-eu.researchsquare.com/files/rs-5190887/v1/ab7cd150d1ee892f3443a53e.docx"}],"financialInterests":"Competing interest reported. CKF owns Microsoft shares. No other potential conflicts of interest were identified by any author.","formattedTitle":"Real-World Evaluation of Artificial Intelligence (AI) Chatbots for Providing Sexual Health Information: A Consensus Study Using Clinical Queries","fulltext":[{"header":"Introduction","content":"\u003cp\u003eArtificial intelligence (AI) has the potential to revolutionise numerous sectors, including healthcare (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e), particularly reshaping health information delivery (\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). AI-powered chatbots could offer a 24/7, non-judgmental and private platform for users to inquire about sensitive or stigmatised health topics, such as sexual health-related issues (\u003cspan additionalcitationids=\"CR5 CR6 CR7 CR8\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eOur study focuses on generative AI, unlike traditional chatbots with pre-programmed responses (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). Generative AIs, like ChatGPT, use natural language processing for human-like conversations (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). They understand context, provide personalised responses, and are more able to engage in dialogue, significantly improving the user experience (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDespite their potential, it is important to consider the possible downsides of AI in the delivery of healthcare information, including the risk of inaccurate or inappropriate advice that may potentially create harms (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). These systems could also inadvertently retain or retrieve sensitive data. These questions raise doubts regarding the extent to which AI-powered chatbots can replace human interaction.\u003c/p\u003e \u003cp\u003eThe Melbourne Sexual Health Centre (MSHC) receives a high volume of phone calls from clients related to sexual health concerns, which places a substantial burden on the centre's nursing staff and limits their availability for direct patient care. To address this issue, MSHC developed a text-based prompt using publicly available information from Australian guidelines for the management of sexually transmitted infections (STIs) (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) and the clinic\u0026rsquo;s website (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis prompt was used to create two customised chatbots through a process known as prompt engineering, which involves crafting specific instructions and providing relevant information to guide AI in generating responses tailored to a particular domain\u0026mdash;in this case, sexual health\u0026mdash;without modifying the underlying AI model. The chatbots developed through this process are \u003cb\u003eAlice\u003c/b\u003e, based on the GPT-3.5 model and implemented on chatbotbuilder.io (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), and \u003cb\u003eAzure\u003c/b\u003e, using the same GPT model on Microsoft Azure (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). In our study, we refer to the refinement of these prompts to further tailor the chatbot responses as prompt tuning. Additionally, we included ChatGPT, developed by OpenAI, without any customisation for comparison.\u003c/p\u003e \u003cp\u003eAlice and Azure chatbots do not have access to or ability to request individual patient records or personal health data. This makes them valuable, low-risk resources for handling common inquiries. Our goal is to evaluate the potential use of these chatbots in handling common questions, allowing nurses to dedicate their efforts to addressing more complex patient needs requiring human attention. Early research indicated that chatbots could deliver accurate information on health topics (\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). However, comparative studies between AI chatbots and human clinicians in the context of sexual health are limited (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). Furthermore, studies examining the effectiveness of different AI chatbots, such as ChatGPT versus prompt-tuned chatbots, are scarce.\u003c/p\u003e \u003cp\u003eThis study aimed to evaluate the capabilities of three different AI chatbots (Alice, Azure, and ChatGPT) compared to clinically trained human staff in responding to sexual health inquiries.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design\u003c/h2\u003e \u003cp\u003eWe conducted this cross-sectional study at the MSHC, Australia's largest public sexual health clinic. We compared the performance of three AI chatbots against experienced sexual health nurses in responding to sexual health inquiries. We used anonymised real-world questions from the callers to the MSHC collected during routine telephone inquiries. In compliance with the Victorian Department of Health guidelines on AI use in healthcare (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e), we did not directly test chatbots with clients. Instead, we designed a method that maintained standard sexual health service delivery (i.e., a telephone conversation with a nurse) while reflecting real-life queries.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eAI Chatbots and Prompt Tuning\u003c/h3\u003e\n\u003cp\u003eWe configured and evaluated three AI chatbots for this study: Alice (Custom GPT-3.5-Turbo on chatbotbuilder.io), Azure (Custom GPT-3.5 on Microsoft Azure), and ChatGPT (standard OpenAI GPT-3.5). We named the chatbots based on their development platforms or origins: \"Alice\" for the chatbot developed on \u003cem\u003echatbotbuilder.io\u003c/em\u003e (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), \"Azure\" for the one implemented on Microsoft Azure, and \"ChatGPT\" for the unmodified OpenAI model. We will use these names consistently throughout this manuscript to refer to these specific chatbot implementations.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e provides a detailed comparison of the chatbots' features and settings. For Alice and Azure, we employed a process known as prompt engineering to create customised chatbots (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e). This involved developing a custom set of instructions (called a \"prompt\") and incorporating a specialised database of information (known as a \"knowledge base\") using publicly available information from the MSHC website (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e) and Australian guidelines for the management of sexually transmitted infections (STIs) (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). This process, which we refer to as \"prompt-tuning\" in this study, involves designing and refining text-based prompts to optimize and tailor the chatbots' responses to sexual health queries. While we didn't modify the underlying AI models, we used these custom prompts and knowledge base to guide the chatbots in providing specialised sexual health information.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Artificial Intelligence (AI) Chatbots' Features and Settings\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeature\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlice \u003c/p\u003e \u003cp\u003e(Custom GPT-3.5-Turbo)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAzure Chatbot \u003c/p\u003e \u003cp\u003e(Custom GPT-3.5)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBase ChatGPT \u003c/p\u003e \u003cp\u003e(OpenAI GPT-3.5)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePlatform\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003echatbotbuilder.io\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMicrosoft Azure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOpenAI\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eModel\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGPT-3.5-Turbo 16K\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGPT-3.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePrompt Tuning\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSpecialisation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePublicly available information \u003c/p\u003e \u003cp\u003e(MSHC website, Australian guidelines for the management of sexually transmitted infections)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePublicly available information \u003c/p\u003e \u003cp\u003e(MSHC website, and Australian guidelines for the management of sexually transmitted infections)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDefault OpenAI training data\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTemperature*\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDefault\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMaximum Tokens**\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDefault\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eData Privacy\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNo patient data access or storage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo patient data access or storage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo patient data access or storage\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eResponse Limitation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUnable to answer questions beyond the provided information\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003cem\u003e*Temperature: A parameter that controls the randomness of the AI's outputs. A lower value (closer to 0) makes the output more focused and deterministic, while a higher value (closer to 1) makes it more diverse and creative.\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003cem\u003e** Maximum Tokens: The maximum number of words or word pieces the AI is allowed to generate in a single response. This limits the length of the AI's answers.\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe development and refinement of Alice, including iterative testing and adjustments, took approximately 4 weeks of dedicated effort from the author (PL). We initially implemented Alice on \u003cem\u003echatbotbuilder.io\u003c/em\u003e (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e) and later applied the same prompt and knowledge base to create Azure on Microsoft Azure (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). Both Alice and Azure used a default temperature setting of 0.5 (which controls the randomness of the AI's responses) and a maximum token limit of 200 for responses (limiting the length of answers). For the Azure chatbot, we configured settings to restrict responses to information derived solely from the provided prompt and knowledge base. This configuration resulted in the chatbot acknowledging its inability to answer questions beyond the scope of the provided information. Such a restrictive setting was not available on the \u003cem\u003echatbotbuilder.io\u003c/em\u003e platform used for Alice. ChatGPT functioned as a control, utilising default OpenAI settings without customised prompt, representing a standard AI chatbot without specific sexual health training. All chatbots operate without access to or storing patient data, ensuring privacy compliance.\u003c/p\u003e\n\u003ch3\u003eCollecting the Sexual Health Queries\u003c/h3\u003e\n\u003cp\u003eBetween January and April 2024, we gathered anonymised questions and responses from calls to the MSHC phone line. To reduce recall bias, sexual health nurses at MSHC documented summaries of clients' questions and responses immediately after each routine telephone consultation. These summaries, recorded through a Microsoft Form, included the clients\u0026rsquo; questions and the nurses\u0026rsquo; answers while carefully excluding all identifying information.\u003c/p\u003e \u003cp\u003eWe gathered a total of 200 question-answer pairs over the four-month period, a sample size determined to ensure a margin of error of approximately\u0026thinsp;\u0026plusmn;\u0026thinsp;7% at a 95% confidence level, balancing statistical precision with feasibility in data collection. The collected questions covered a range of sexual health topics, including STI symptoms and testing, contraception methods, clinic services and hours, and general sexual health advice.\u003c/p\u003e\n\u003ch3\u003ePreparing and Processing the Data\u003c/h3\u003e\n\u003cp\u003ePrior to analysis, two researchers (PL and NS) verified that all summaries were free of identifiable information. We then input the summarised questions into each of the three AI chatbots: Alice, Azure, and ChatGPT. To prevent context bias, we entered each question into a new session for each chatbot. This process resulted in a total of 600 responses: 200 from each of the three chatbots, in addition to the 200 answers provided by nurses.\u003c/p\u003e\n\u003ch3\u003eExpert Evaluation and Consensus Process\u003c/h3\u003e\n\u003cp\u003eAfter data collection, we assembled a panel of three experts, who were both sexual health physicians and researchers, to evaluate the quality and accuracy of the responses between June and July 2024. The reviewers had between 7 and 25 years of experience in sexual health medicine (EA, KH, CKF) with extensive expertise in caring for patients seen at the study clinic. Prior to the evaluation, the research team and reviewers collaboratively defined and clarified the meaning of five key outcome measures: guidance, accuracy, safety, ease of access, and provision of only necessary information. This ensured a consistent interpretation of the criteria throughout the evaluation process. (See Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOutcome Indicator Definition\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOutcome Measure\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDefinition\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOverall Correctness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses the overall accuracy and appropriateness of the response, considering both factual correctness and suitability for the given question.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGuidance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses whether the patient will take the appropriate action or make the right decision after reading the response.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses the correctness of the information provided in the response.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSafety\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses the potential risk of harm to the patient if they follow the advice given in the response. This includes potential conflict with health care providers from wrong advice.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEase of Understanding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses the clarity and readability of the response for a general audience.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProvision of Necessary Information Only\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAssesses whether the response provides concise, relevant information without including unnecessary details that could deter the patient from using the chatbot.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003eNote: The outcome measures are considered independent of each other. For example, a response can have excellent ease of understanding while inaccurate.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ePL developed a Qualtrics survey and conducted a pilot test with five questions and respective answers. This pilot allowed reviewers to familiarise themselves with the rating process, apply the agreed-upon definitions, and estimate the time required for the full review. In the survey, we labelled nurses' summaries as 'Nurses' and assigned anonymous identifiers to the three chatbot responses. To minimise bias, we designed a blinded review process. For each question set, we consistently presented the nurse's summary first due to its distinctive appearance (i.e., in note format rather than verbatim). The three chatbot responses followed in a randomised order, enabling blinded comparisons between the AI chatbots.\u003c/p\u003e \u003cp\u003eFollowing the pilot test, the team held a consensus meeting to discuss score discrepancies, explain reasoning based on established definitions, and reach a unified judgment. Each reviewer then independently evaluated the remaining 195 questions and answers using the agreed-upon outcome measures and rating scale.\u003c/p\u003e \u003cp\u003eFollowing individual evaluations, we conducted a final consensus process. We categorised responses into binary classifications for correctness and the five outcome measures. We identified cases where two reviewers agreed, but the third differed. In a consensus meeting, reviewers discussed these discrepancies and worked towards a unified assessment. This process ensured evaluation consistency while preserving the integrity of initial judgments.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eWe used STATA (version 17, StataCorp) for data analysis in this study. To describe response lengths, we calculated the median and interquartile range (IQR) of word count for each chatbot and nurses\u0026rsquo; responses.\u003c/p\u003e \u003cp\u003eFor all responses, we categorised overall correctness into \"correct\" (combining \"correct\" and \"mostly correct\" ratings) and \"incorrect\" (combining \"partially correct\" and \"incorrect\" ratings). For the five outcome measures (guidance, accuracy, safety, ease of access, and provision of necessary information), we classified responses as \"acceptable or better\" (including \"acceptable\", \"good\", and \"excellent\" ratings) or \"unacceptable\" (including \"poor\" and \"very poor\" ratings).\u003c/p\u003e \u003cp\u003eWe then calculated the proportion of correct or acceptable responses for each chatbot and nurses, using questions with correct nurse responses as the benchmark. We compared these proportions using chi-square tests, considering p-values less than 0.05 as statistically significant. We calculated 95% confidence intervals for all proportions.\u003c/p\u003e \u003cp\u003eWe conducted subgroup analyses by stratifying questions into \"General sexual health questions\" and \"Clinic-specific questions,\" repeating our performance comparisons for each subgroup. To account for the differences in chatbot configurations, particularly Azure's restricted response settings, we performed a sensitivity analysis to assess the impact of Azure's restricted configuration on our overall findings and ensure a fair comparison across all chatbots. In this analysis, we reran our main analyses after excluding questions that the Azure chatbot could not answer due to its limitation to the provided prompt and knowledge base.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eEthical Considerations\u003c/h3\u003e\n\u003cp\u003e This study received ethical approval from the Alfred Hospital Ethics Committee in Melbourne, Australia (project number: 555/23). The committee waived the need for informed consent on the basis that the study used de-identified, routinely collected clinical data. All research procedures adhered to the committee's advice and Australian ethical standards for clinical research. No identifiable information was collected or stored during the study process.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eWe analysed a total of 195 questions, following the exclusion of five questions used in pilot testing. Each question had four responses: one from a nurse and one from each of the three AI chatbots, resulting in a total of 780 responses.\u003c/p\u003e\n\u003cp\u003eThe median length of client questions was 14 words (IQR, 10\u0026ndash;19). Nurses\u0026apos; responses were concise as they were written in the form of a summary note, with a median length of 25 words (IQR, 13\u0026ndash;39). In contrast, AI-generated responses were substantially longer. Alice produced responses with a median length of 105 words (IQR, 81\u0026ndash;147), ChatGPT a median of 137 words (IQR, 92\u0026ndash;224) and Azure a median length of 83 words (IQR, 50\u0026ndash;122) (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eOf the 195 questions analysed, nurses provided correct responses to 192 (98.5%) queries. It\u0026apos;s important to note that for nurse responses, we assessed only overall correctness, whereas for chatbots, we evaluated both overall correctness and five additional outcome measures. Therefore, comparisons between nurses and chatbots are limited to overall correctness, while the five outcome measures are compared only among the three chatbots. We based our analysis of chatbot performance on these 192 correctly answered questions.\u003c/p\u003e\n\u003cp\u003eAlice demonstrated the highest overall correctness at 85.2% (95% CI, 82.1%-88.0%), followed by Azure at 69.3% (95% CI, 65.3%-73.0%), and ChatGPT at 64.8% (95% CI, 60.7%-68.7%). The difference in performance between the three chatbots was statistically significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). Regarding the five outcome measures, Alice consistently outperformed the other chatbots in guidance (90.1%; 95% CI, 87.4%-92.4%) and accuracy (87.8%; 95% CI, 84.9%-90.4%). Azure showed strength in safety (97.9%; 95% CI, 96.4%-98.9%) and ease of access (95.1%; 95% CI, 93.1%-96.7%), outperforming Alice and ChatGPT in these areas and these findings are statistically significant with p-values less than 0.05. ChatGPT consistently showed lower scores across all measures compared to Alice and Azure (Table \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). The distribution of individual scores for each outcome measure is illustrated in Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eSubgroup analysis\u003c/p\u003e\n\u003cp\u003eWe conducted a subgroup analysis by dividing the questions into clinic-related questions (n\u0026thinsp;=\u0026thinsp;78) and general sexual health questions (n\u0026thinsp;=\u0026thinsp;114). (Table S1)\u003c/p\u003e\n\u003cp\u003eFor clinic-related questions, Alice demonstrated the highest overall correctness at 78.2% (95% CI, 72.4%-83.3%), followed by Azure (69.2%; 95% CI, 62.9%-75.1%) and ChatGPT (39.3%; 95% CI, 33.0%-45.9%). Across all outcome measures for clinic-related questions, Alice and Azure consistently outperformed ChatGPT. Notably, all chatbots achieved high scores in safety, with Azure the highest (98.7%; 95% CI, 96.3%-99.7%), followed by Alice (94.9%; 95% CI, 91.2%-97.3%), and ChatGPT (91.0%; 95% CI, 86.6%-94.4%).\u003c/p\u003e\n\u003cp\u003eFor general sexual health questions, Alice demonstrated the highest overall correctness at 90.1% (95% CI, 86.4%-93.0%), followed by ChatGPT at 82.2% (95% CI, 77.7%-86.1%) and Azure at 69.3% (95% CI, 64.1%-74.1%). Across all outcome measures for general sexual health questions, Alice consistently outperformed the other chatbots. Notably, all chatbots achieved high scores in safety, with Alice and Azure both at 97.4% (95% CI, 95.1%-98.8%), followed closely by ChatGPT at 96.5% (95% CI, 94.0%-98.2%).\u003c/p\u003e\n\u003cp\u003eSensitivity analysis\u003c/p\u003e\n\u003cp\u003eWe conducted a sensitivity analysis after removing questions that Azure was unable to answer due to its restricted configuration. This analysis included 168 questions, resulting in 504 total responses across the three chatbots. (Table S2)\u003c/p\u003e\n\u003cp\u003eIn this sensitivity analysis, the performance gap between Alice and Azure narrowed considerably. Alice\u0026apos;s overall correctness was 83.9% (95% CI, 80.4%-87.0%) compared to Azure\u0026apos;s 79.2% (95% CI, 75.4%-82.6%) (p\u0026thinsp;=\u0026thinsp;0.051). Both chatbots showed comparable results in safety (Alice 95.8%, Azure 97.6%, p\u0026thinsp;=\u0026thinsp;0.1) and only necessary information given (Alice 91.1%, Azure 94.4%, p\u0026thinsp;=\u0026thinsp;0.6). Alice had higher guidance (88.7% vs 82.7%, p\u0026thinsp;=\u0026thinsp;0.007) and accuracy (86.3% vs 81.0%, p\u0026thinsp;=\u0026thinsp;0.02), while Azure had higher ease of access (94.4% vs 91.1%, p\u0026thinsp;=\u0026thinsp;0.04).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur study revealed significant differences in performance among the three AI chatbots - Alice, Azure, and ChatGPT- in responding to sexual health questions based on actual inquiries to a sexual health clinic telephone line. The prompt-tuned chatbots, Alice and Azure, consistently outperformed the standard ChatGPT across all measures, with Alice demonstrating the highest overall correctness at 85.2%. This superior performance extended across various outcome measures, including guidance, accuracy, and provision of necessary information. Notably, all chatbots achieved high safety scores, particularly Azure reaching 97.9% in this important measure. Our subgroup analysis further highlighted the strengths of prompt-tuned chatbots, particularly in handling clinic-specific queries where Alice and Azure significantly outperformed ChatGPT. These findings align with recent research by Koh et al., who found that ChatGPT could provide helpful and accurate information regarding STIs, although they noted that the advice lacked specificity and required human oversight (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). Our study extends these findings by demonstrating the potential benefits of prompt-tuning in enhancing AI performance for providing sexual health information. Our findings underscore the importance of continuous improvement in prompt engineering and knowledge base development for AI chatbots in healthcare applications (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e). The significance of these findings lies in demonstrating that AI chatbots, especially those with domain-specific training, can provide accurate and safe sexual health information across a range of query types (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eOur study found that prompt-tuned AI chatbots significantly outperformed the standard ChatGPT in providing sexual health information. This aligns with previous research showing the benefits of domain-specific training for AI in healthcare applications (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). The superior performance of Alice and Azure, particularly in areas like guidance and accuracy, suggests that tailored knowledge bases and specialised prompts are crucial for AI's effectiveness in healthcare contexts. However, we found that all AI chatbots, including the prompt-tuned chatbots, are susceptible to hallucinations - the generation of false or irrelevant information that is not grounded in the provided data or knowledge base - at some point. This issue was particularly prominent with ChatGPT, which again is in line with other studies (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). This limitation highlights the need for ongoing refinement and regular updates of AI models.\u003c/p\u003e \u003cp\u003eThe high safety scores achieved by the chatbots, particularly Azure at 98%, is an important finding with significant implications for the potential integration of AI in healthcare settings. In this study, the reviewers paid particular attention to safety scores because previous studies have raised concerns about the safety of AI-generated health advice (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e), making our results particularly noteworthy. Koh et al. also found that ChatGPT provided generally safe advice for STI-related queries, although they emphasised the need for human physician involvement (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). The fact that prompt-tuned chatbots achieved such high safety scores suggests that careful design and training might mitigate many risks associated with AI in healthcare. However, these promising results do not guarantee absolute safety in real-world applications, nor do they negate the need for human oversight. Further studies, including real-world implementations under controlled conditions, are needed to validate these findings and assess the true safety of AI chatbots in healthcare settings. It is important to note that even advice from clinical staff is not 100% safe, highlighting the complexity of healthcare information delivery.\u003c/p\u003e \u003cp\u003eThe integration of AI chatbots for providing sexual health information raises important ethical considerations. Privacy and data security are paramount, given the sensitive nature of sexual health information (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e). Real-world implementation would require robust safeguards against data breaches and unauthorised access. Balancing AI efficiency with the irreplaceable human element in healthcare is crucial. Clear guidelines must be established for when AI should defer to human expertise, and patients should be made aware when they are interacting with an AI system rather than a human healthcare provider.\u003c/p\u003e \u003cp\u003eThis study offers several key strengths in the methodology and design. Foremost is our use of real-world sexual health queries and responses from Australia's largest sexual health clinic, ensuring high ecological validity and relevance to actual clinical practice. We implemented a multi-step, consensus-based review process involving expert clinicians, enhancing our evaluations' reliability and depth. Our iterative consensus-based approach differs from traditional methods that often rely on averaged scores or majority decisions. By actively addressing discrepancies and seeking unanimous agreement, we aimed to enhance the reliability and depth of our evaluations. This process allowed us to capture nuanced insights that might be lost in purely quantitative assessments and establish more definitive benchmarks for AI performance in providing sexual health information. While time-intensive, this method allows for a comprehensive exploration of evaluation criteria, potentially resulting in more robust and clinically relevant outcomes. Importantly, our study goes beyond evaluating just the ChatGPT; we included two prompt-tuned versions (Alice and Azure), providing insights into the potential of customised AI chatbots for specialised healthcare domains. This comparison of base and prompt-tuned chatbots offers valuable data on the effectiveness of domain-specific AI training. Additionally, our comprehensive assessment across five key outcome measures, including the critical aspect of safety, provides a multifaceted analysis of AI performance in a sensitive healthcare context.\u003c/p\u003e \u003cp\u003eOur study has several limitations. First, while our sample size of 195 questions was sufficient for our exploratory aims, a larger sample would improve the precision. Second, the data collection method, using nurse-summarised responses rather than verbatim transcripts, was chosen for feasibility and to protect patient privacy. While potentially introducing some level of recall and observer bias, this approach aligns with our goal of using nurse responses as a benchmark rather than for direct comparison with chatbots. Moreover, as the summaries were written by the nurses themselves, they are likely to accurately reflect the key points of the interactions. Third, the study's focus on a single sexual health clinic, despite being the largest sexual health centre in Australia, may limit the diversity of queries represented. Additionally, the different operational settings of the chatbots - with Azure's responses limited to its prompt and knowledge base, while Alice had no such restrictions - could introduce performance bias. We addressed this through sensitivity analysis, comparing performance with and without Azure's limited responses. Lastly, the rapid evolution of AI technology means that chatbot performance may have changed since data collection, potentially affecting the long-term applicability of our results.\u003c/p\u003e \u003cp\u003eThe findings of this study have significant implications for the future of providing sexual health and broader healthcare information. By handling routine inquiries, AI chatbots could potentially reduce the workload on healthcare professionals, allowing them to focus on the delivery of clinical care requiring human expertise. The superior performance of our prompt-tuned chatbots underscores the importance of domain-specific training in AI applications for healthcare. This implies that customised AI chatbots could be developed for various medical specialities, enhancing the quality and accessibility of health information across diverse fields. Moreover, the 24/7 availability and instant responses of AI chatbots could improve access to sexual health information, potentially enabling earlier detection and treatment of STIs. The high safety scores across two prompt-tuned chatbots are particularly encouraging, suggesting that AI could be a reliable source of health information with proper development and oversight.\u003c/p\u003e \u003cp\u003eFuture research should expand the scope of AI chatbot evaluations in healthcare, using larger, more diverse datasets across various demographics and clinical settings. Longitudinal studies are needed to assess chatbot performance over time as AI technologies evolve. It is also important to investigate real-time patient interactions and integrate AI into hybrid care models, where AI assists human clinicians.\u003c/p\u003e \u003cp\u003eIn conclusion, our study demonstrates the potential of prompt-tuned AI chatbots in providing accurate and safe sexual health information. While promising, these findings also highlight the need for continued research, development, and careful implementation to fully realise the benefits of AI in healthcare settings.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eWe thank the nurses at the Melbourne Sexual Health Centre for their invaluable contribution to collecting data amidst their busy phone room duties. We appreciate Mark Chung\u0026apos;s insightful suggestions and his assistance with the data collection. We acknowledge Monash University for providing PL with a PhD scholarship.\u003c/p\u003e\n\u003ch2\u003eAuthor Contributions\u003c/h2\u003e\n\u003cp\u003eCKF, PL and LZ conceived the study. CKF provided overall supervision and was integrally involved in all stages of the research. CKF guided the study design, chatbot development, data collection process, and analysis while also serving as a key reviewer in the blinded evaluation of the chatbots. PL led the chatbot development, designed the study, prepared the questionnaire, created the Qualtrics survey, conducted data analysis, and wrote the initial manuscript draft. NS contributed to Azure chatbot development and data collection. DL assisted with chatbot development. RF, SD, and CR assisted with data collection. CKF, EA, and KH contributed as reviewers, and JO contributed to reviewer meetings. All authors contributed to the manuscript revision and approved the final version.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eCKF is supported by a National Health and Medical Research Council (NHMRC) Leadership Investigator Grant (GNT1172900). JO and EPF are each supported by the\u0026nbsp;NHMRC Emerging Leadership Investigator Grant (GNT1193955 and\u0026nbsp;GNT1172873, respectively).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003ePotential Conflicts of Interest\u003c/h2\u003e\n\u003cp\u003eCKF owns Microsoft shares. No other potential conflicts of interest were identified by any author.\u003c/p\u003e\n\u003ch2\u003eData Availability\u003c/h2\u003e\n\u003cp\u003eData available upon reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAlowais SA, Alghamdi SS, Alsuhebany N, Alqahtani T, Alshaya AI, Almohareb SN, et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Medical Education. 2023;23(1):689.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang J, Oh YJ, Lange P, Yu Z, Fukuoka Y. Artificial Intelligence Chatbot Behavior Change Model for Designing Artificial Intelligence Chatbots to Promote Physical Activity and a Healthy Diet: Viewpoint. J Med Internet Res. 2020;22(9):e22845.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiao Z, Liao V, Zhou M, Grandison T, Li Y. Powering an AI Chatbot with Expert Sourcing to Support Credible Health Information Access2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhawaja Z, B\u0026eacute;lisle-Pipon JC. Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots. Front Digit Health. 2023;5:1278186.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Gupta S, Singhal A, Muttreja P, Singh S, Sharma P, et al. An Artificial Intelligence Chatbot for Young People\u0026rsquo;s Sexual and Reproductive Health in India (SnehAI): Instrumental Case Study. Journal of Medical Internet Research. 2022;24(1):e29969.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNadarzynski T, Puentes V, Pawlak I, Mendes T, Montgomery I, Bayley J, et al. Barriers and facilitators to engagement with artificial intelligence (AI)-based chatbots for sexual and reproductive health advice: a qualitative analysis. Sexual Health. 2021;18(5):385\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMills R, Mangone ER, Lesh N, Mohan D, Baraitser P. Chatbots to Improve Sexual and Reproductive Health: Realist Synthesis. J Med Internet Res. 2023;25:e46761.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiklosik A, Evans N, Qureshi A. The Use of Chatbots in Digital Business Transformation: A Systematic Literature Review. IEEE Access. 2021;9:106530\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan H, Han B, Gao W, Li W. How AI chatbots have reshaped the frontline interface in China: examining the role of sales\u0026ndash;service ambidexterity and the personalization\u0026ndash;privacy paradox. International Journal of Emerging Markets. 2022;ahead-of-print.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStokel-Walker C, Van Noorden R. What ChatGPT and generative AI mean for science. Nature. 2023;614(7947):214\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeng J, Lin Y. The Benefits and Challenges of ChatGPT: An Overview. Frontiers in Computing and Intelligent Systems. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePatel AS. Docs get clever with ChatGPT. Medscape. February 3, 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWasson EJ, Driver K, Hughes M, Bailey J. Sexual reproductive health chatbots: should we be so quick to throw artificial intelligence out with the bathwater? BMJ Sexual \u0026amp; Reproductive Health. 2021;47(1):73-.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrown JEH, Halpern J. AI chatbots cannot replace human interactions in the pursuit of more inclusive mental healthcare. SSM - Mental Health. 2021;1:100017.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOng JJ, Bourne C, Dean JA, Ryder N, Cornelisse VJ, Murray S, et al. Australian sexually transmitted infection (STI) management guidelines for use in primary care 2022 update. Sex Health. 2023;20(1):1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMelbourne Sexual Health Centre. [Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.mshc.org.au/\u003c/span\u003e\u003cspan address=\"https://www.mshc.org.au/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAI CB. 2024 [Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.chatbotbuilder.ai/\u003c/span\u003e\u003cspan address=\"https://www.chatbotbuilder.ai/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMicrosoft. Azure AI Bot Service 2024 [Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://azure.microsoft.com/en-au/products/ai-services/ai-bot-service\u003c/span\u003e\u003cspan address=\"https://azure.microsoft.com/en-au/products/ai-services/ai-bot-service\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNadarzynski T, Bayley J, Llewellyn C, Kidsley S, Graham CA. Acceptability of artificial intelligence (AI)-enabled chatbots, video consultations and live webchats as online platforms for sexual health advice. BMJ Sex Reprod Health. 2020;46(3):210\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePotapenko I, Boberg-Ans LC, Michael, Klefter ON, Van Dijk EHC, Subhi Y. Artificial intelligence‐based chatbot patient information on common retinal diseases using\u0026thinsp;\u0026lt;\u0026thinsp;scp\u0026thinsp;\u0026gt;\u0026thinsp;ChatGPT\u0026lt;/scp\u0026gt;. Acta Ophthalmologica. 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSecinaro S, Calandra D, Secinaro A, Muthurangu V, Biancone P. The role of artificial intelligence in healthcare: a structured literature review. BMC Medical Informatics and Decision Making. 2021;21(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAyers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023;183(6):589\u0026ndash;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVictoria Department of Health. Health service use of unregulated Artificial Intelligence (AI) 2023 [Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.safercare.vic.gov.au/sites/default/files/2023-07/Advisory%20-%20ChatGPT%20and%20Generative%20AI%20July%202023%20FINAL.pdf\u003c/span\u003e\u003cspan address=\"https://www.safercare.vic.gov.au/sites/default/files/2023-07/Advisory%20-%20ChatGPT%20and%20Generative%20AI%20July%202023%20FINAL.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLund B. The prompt engineering librarian. Library Hi Tech News. 2023;40(8):6\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoh MCY, Ngiam JN, Tambyah PA, Archuleta S. ChatGPT as a tool to improve access to knowledge on sexually transmitted infections. Sex Transm Infect. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKozaily E, Geagea M, Akdogan ER, Atkins J, Elshazly MB, Guglin M, et al. Accuracy and consistency of online large language model-based artificial intelligence chat platforms in answering patients' questions about heart failure. Int J Cardiol. 2024;408:132115.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMart\u0026iacute;nez-Ezquerro JD. Response to: Impact of ChatGPT and Artificial Intelligence in the Contemporary Medical Landscape. Archives of Medical Research. 2023;54(5):102838.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbbasian M, Khatibi E, Azimi I, Oniani D, Shakeri Hossein Abad Z, Thieme A, et al. Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI. NPJ Digit Med. 2024;7(1):82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBirkun AA, Gautam A. Large Language Model (LLM)-Powered Chatbots Fail to Generate Guideline-Consistent Content on Resuscitation and May Provide Potentially Harmful Advice. Prehospital and Disaster Medicine. 2023;38(6):757\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrabb D, Lamparth M, Vasan N. Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation. medRxiv. 2024:2024.04.07.24305462.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKing AJ, Latt, P. M., Zhang, L., Soe, N. N., Temple-Smith, M., Maddaford, K., Fairley, C. K., Chow, E. P. F., \u0026amp; Phillips, T. R. User experience of an AI application for predicting risk of sexually transmitted infections: A qualitative study. 2024.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable\u0026nbsp;1. Comparison of Artificial Intelligence (AI) Chatbots\u0026apos; Features and Settings\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFeature\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAlice\u0026nbsp;\u003cbr\u003e\u0026nbsp;(Custom GPT-3.5-Turbo)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAzure Chatbot\u0026nbsp;\u003cbr\u003e\u0026nbsp;(Custom GPT-3.5)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBase ChatGPT\u0026nbsp;\u003cbr\u003e\u0026nbsp;(OpenAI GPT-3.5)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePlatform\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003echatbotbuilder.io\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003eMicrosoft Azure\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eOpenAI\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003eGPT-3.5-Turbo 16K\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003eGPT-3.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eGPT-3.5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrompt Tuning\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpecialisation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003ePublicly available information\u0026nbsp;\u003cbr\u003e\u0026nbsp;(MSHC website, Australian guidelines for the management of sexually transmitted infections)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003ePublicly available information\u0026nbsp;\u003cbr\u003e\u0026nbsp;(MSHC website, and Australian guidelines for the management of sexually transmitted infections)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eDefault OpenAI training data\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTemperature*\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003e0.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003e0.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eDefault\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMaximum Tokens**\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003e200\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003e200\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eDefault\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eData Privacy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003eNo patient data access or storage\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003eNo patient data access or storage\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eNo patient data access or storage\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 18.3673%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eResponse Limitation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 28.5714%;\"\u003e\n \u003cp\u003eNone\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.6122%;\"\u003e\n \u003cp\u003eUnable to answer questions beyond the provided information\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 22.449%;\"\u003e\n \u003cp\u003eNone\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cem\u003e*Temperature: A parameter that controls the randomness of the AI\u0026apos;s outputs. A lower value (closer to 0) makes the output more focused and deterministic, while a higher value (closer to 1) makes it more diverse and creative.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e**\u003c/em\u003e \u003cem\u003eMaximum Tokens: The maximum number of words or word pieces the AI is allowed to generate in a single response. This limits the length of the AI\u0026apos;s answers.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eTable\u0026nbsp;2. Outcome Indicator Definition\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eOutcome Measure\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDefinition\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eOverall Correctness\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses the overall accuracy and appropriateness of the response, considering both factual correctness and suitability for the given question.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eGuidance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses whether the patient will take the appropriate action or make the right decision after reading the response.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eAccuracy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses the correctness of the information provided in the response.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eSafety\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses the potential risk of harm to the patient if they follow the advice given in the response. This includes potential conflict with health care providers from wrong advice.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eEase of Understanding\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses the clarity and readability of the response for a general audience.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eProvision of Necessary Information Only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAssesses whether the response provides concise, relevant information without including unnecessary details that could deter the patient from using the chatbot.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNote: The outcome measures are considered independent of each other. For example, a response can have excellent ease of understanding while inaccurate.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eTable\u0026nbsp;\u003c/em\u003e\u003cem\u003e3\u003c/em\u003e\u003cem\u003e. Overall Performance of the Three AI Chatbots in Responding to Sexual Health Queries\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003eOutcome Measure\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"3\"\u003e\n \u003cp\u003e\u003cstrong\u003eThe proportion of Overall Correctness by Chatbots\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(N =576) [192 questions]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026nbsp;P-value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 142px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAlice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 130px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 125px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAzure\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eOverall Correctness\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e85.2%\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(82.1%-88.0%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e64.8%\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(60.7%-68.7%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e69.3%\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(65.3%-73.0%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026lt;0.0001\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eOutcome Measures\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"4\"\u003e\n \u003cp\u003e\u003cstrong\u003eThe proportion of Acceptable and Beyond Responses by Chatbots\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e(N =576) [192 questions]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eGuidance\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e90.1%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(87.4%-92.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e66.1%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(62.1%-70.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e84.9%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(81.7%-87.7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e87.8%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(84.9%-90.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e64.9%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(60.9%-68.8%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e70.8%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(66.9%-74.5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eSafety\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e96.4%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(94.5%-97.7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e94.3%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(92.0%-96.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e97.9%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(96.4%-98.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0.005\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eEase of Access\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e91.7%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(89.1%-93.8%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e77.6%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(74.0%-80.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e95.1%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(93.1%-96.7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eOnly Necessary Information Given\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e91.3%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(88.7%-93.5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e69.8%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(65.9%-73.5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e92.7%\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e(90.3%-94.7%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u0026lt;0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5190887/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5190887/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eIntroduction\u003c/h2\u003e \u003cp\u003eArtificial Intelligence (AI) chatbots could potentially provide information on sensitive topics, including sexual health, to the public. However, their performance compared to human clinicians and across different AI chatbots, particularly in the field of sexual health, remains understudied. This study evaluated the performance of three AI chatbots - two prompt-tuned (Alice and Azure) and one standard chatbot (ChatGPT by OpenAI) - in providing sexual health information, compared to human clinicians.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe analysed 195 anonymised sexual health questions received by the Melbourne Sexual Health Centre phone line. A panel of experts in a blinded order using a consensus-based approach evaluated responses to these questions from nurses and the three AI chatbots. Performance was assessed based on overall correctness and five specific measures: guidance, accuracy, safety, ease of access, and provision of necessary information. We conducted subgroup analyses for clinic-specific (e.g., opening hours) and general sexual health questions and a sensitivity analysis excluding questions that Azure could not answer.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eAlice demonstrated the highest overall correctness (85.2%; 95% confidence interval (CI), 82.1%-88.0%), followed by Azure (69.3%; 95% CI, 65.3%-73.0%) and ChatGPT (64.8%; 95% CI, 60.7%-68.7%). Prompt-tuned chatbots outperformed the base ChatGPT across all measures. Azure achieved the highest safety score (97.9%; 95% CI, 96.4%-98.9%), indicating the lowest risk of providing potentially harmful advice. In subgroup analysis, all chatbots performed better on general sexual health questions compared to clinic-specific queries. Sensitivity analysis showed a narrower performance gap between Alice and Azure when excluding questions Azure could not answer.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003ePrompt-tuned AI chatbots demonstrated superior performance in providing sexual health information compared to base ChatGPT, with high safety scores particularly noteworthy. However, all AI chatbots showed susceptibility to generating incorrect information. These findings suggest the potential for AI chatbots as adjuncts to human healthcare providers for providing sexual health information while highlighting the need for continued refinement and human oversight. Future research should focus on larger-scale evaluations and real-world implementations.\u003c/p\u003e","manuscriptTitle":"Real-World Evaluation of Artificial Intelligence (AI) Chatbots for Providing Sexual Health Information: A Consensus Study Using Clinical Queries","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-10-09 05:38:25","doi":"10.21203/rs.3.rs-5190887/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0cab567d-584d-40d7-b3cd-f4b9a868c4c7","owner":[],"postedDate":"October 9th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":38648164,"name":"Health sciences/Health care"},{"id":38648165,"name":"Health sciences/Medical research"},{"id":38648166,"name":"Health sciences/Risk factors"},{"id":38648167,"name":"Health sciences/Signs and symptoms"}],"tags":[],"updatedAt":"2024-10-09T05:38:25+00:00","versionOfRecord":[],"versionCreatedAt":"2024-10-09 05:38:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5190887","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5190887","identity":"rs-5190887","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.