Enhancing AI Chatbot Responses in Healthcare: The SMART Prompt Structure in Head and Neck Surgery | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Enhancing AI Chatbot Responses in Healthcare: The SMART Prompt Structure in Head and Neck Surgery Luigi Angelo Vaira, Jerome R. Lechien, Vincenzo Abbate, Guido Gabriele, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4953716/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Objective. To evaluate the impact of prompt construction on the quality of AI chatbot responses in the context of head and neck surgery. Study design. Observational and evaluative study. Setting. International collaboration involving 16 researchers from 11 European centers specializing in head and neck surgery. Methods. A total of 24 questions, divided into clinical scenarios, theoretical questions, and patient inquiries, were developed. These questions were inputted into ChatGPT-4o both with and without the use of a structured prompt format, known as SMART (Seeker, Mission, AI Role, Register, Targeted Question). The AI-generated responses were evaluated by experienced head and neck surgeons using the QAMAI instrument, which assesses accuracy, clarity, relevance, completeness, source quality, and usefulness. Results. The responses generated using the SMART prompt scored significantly higher across all QAMAI dimensions compared to those without contextualized prompts. Median QAMAI scores for SMART prompts were 27.5 (IQR 25–29) versus 24 (IQR 21.8–25) for unstructured prompts (p < 0.001). Clinical scenarios and patient inquiries showed the most significant improvements, while theoretical questions also benefited but to a lesser extent. The AI's source quality improved notably with the SMART prompt, particularly in theoretical questions. Conclusions. The study suggests that the structured SMART prompt format significantly enhances the quality of AI chatbot responses in head and neck surgery. This approach improves the accuracy, relevance, and completeness of AI-generated information, underscoring the importance of well-constructed prompts in clinical applications. Further research is warranted to explore the applicability of SMART prompts across different medical specialties and AI platforms. Artificial Intelligence and Machine Learning Otorhinolaryngology ChatGPT artificial intelligence AI prompt engineering maxillofacial surgery otorhinolaryngology 1. INTRODUCTION In recent years, the integration of Artificial Intelligence (AI) within the healthcare sector has rapidly expanded, offering unprecedented opportunities to enhance patient care, streamline administrative processes, and support clinical decision-making 1 – 3 . Among the various AI applications, conversational agents, commonly known as chatbots, have gained significant attention for their potential to interact with users in a natural language format. These AI-driven systems are being increasingly utilized for a wide range of purposes, including symptom checking, patient triage, mental health support, and health information dissemination 4 – 8 . The effectiveness of AI chatbots in healthcare, however, largely depends on the quality of the interactions they facilitate 9 , 10 . A key factor influencing these interactions is the way users formulate their prompts or queries 11 , 12 . Unlike traditional search engines, where users might rely on keyword-based queries, interactions with AI chatbots often involve more complex, context-rich language, which can significantly affect the responses generated by these systems. This is particularly crucial in healthcare settings, where the accuracy, clarity, and relevance of information can have profound implications on patient outcomes. Providing adequate context to an AI chatbot is crucial for obtaining accurate and contextually appropriate responses. The quality of AI-generated responses is significantly influenced by the amount and relevance of the information provided in the prompt 13 . When a prompt lacks sufficient context, the AI is forced to make broader assumptions, which can lead to generalized or less accurate responses. Conversely, when a prompt includes detailed and specific information, such as patient history, symptoms, or the specific surgical context, the AI can tailor its responses more precisely to the needs of the user. Despite the growing use of AI chatbots, there is limited number of research exploring how the structuring of prompts influences the quality of responses in healthcare-related contexts 11 , 12 . Understanding this relationship is essential for optimizing the design and deployment of AI systems in healthcare, ensuring they provide reliable and useful information to users. Moreover, as AI continues to evolve, developing best practices for prompt formulation could enhance the overall user experience and effectiveness of AI-driven healthcare services. This study aims to investigate the impact of prompt construction on the quality of AI chatbot responses specifically within the context of head and neck surgery. 2. MATERIALS AND METHODS This study was conducted as part of an international collaborative research project involving young researchers from the International Federation of Otorhinolaryngology Societies and the Italian Society of Maxillofacial Surgery. The consortium was established in February 2023, with the purpose of exploring the potential applications, assessing the reliability, and identifying possible risks associated with artificial intelligence (AI) platforms in the field of head and neck surgery. Sixteen researchers from 11 European centers participated in this study. The requirement for an ethical review and approval was waived because the study did not include any analysis of humans or animals. 2.1 Prompt development For the specific purpose of developing the prompt to be used in this study, a multidisciplinary team was constituted. This team included two head and neck surgeons, a linguist, and a computer engineer. Their objective was to create a prompt that would effectively guide the AI in providing accurate and contextually appropriate responses. The prompt format was developed through multiple rounds of testing and refinement. Initial versions of the prompt were tested using various clinical scenarios provided by the head and neck surgeons. Feedback was collected and analyzed to identify any areas where the AI responses were suboptimal. Adjustments were made to the prompt format to improve clarity, specificity, and relevance. The final prompt structure was validated by applying it across different clinical cases to ensure its effectiveness in generating high-quality AI responses. This process resulted in the creation of a prompt format identified by the acronym SMART, which stands for Seeker, Mission, AI Role, Register, and Targeted Question. Seeker : Refers to the identity and perspective of the user inquiring (i.e. head and neck surgeon, general practitioner, medical student, patient). The prompt must communicate the seeker’s role and the clinical context to ensure the AI understands the expertise level and specific needs of the user. Mission : Defines the purpose of the inquiry. This involves clearly stating the objective or problem that the AI is expected to address. The mission statement ensures the AI focuses on the most relevant information. AI Role : Describes the expected role of the AI in the interaction, which may range from providing information to offering clinical advice. Clearly defining the AI’s role helps set expectations for the type and depth of the response. Register : Involves the tone and style of language to be used in the AI’s response. The prompt should guide the AI to use language that is appropriate for the target audience, ensuring clarity and avoiding ambiguity. Targeted Question : Refers to the specific question or query posed to the AI. This question should be direct and focused, designed to elicit a detailed and accurate response. 2.2 Study design A separate group of three head and neck surgeons, distinct from those involved in the prompt development, was tasked with creating a set of 24 questions. These questions were divided into three groups of eight questions each, covering different types of inquiries: clinical scenarios, theoretical questions, and patient inquiries. The questions encompassed three main areas of head and neck surgery: oncology, sinus and nasal surgery, and trauma. The questions were carefully reviewed to eliminate any potential ambiguities, ensuring clarity and precision. The questions were formulated in a question format suitable for direct input into the chatbot. Additionally, all questions required the chatbot to provide bibliographic references for the information it utilized in generating responses. The same research group also prepared the contextual information to be included in the SMART prompt [Table 1 ]. The "Targeted Question" section of the SMART prompt was then populated with the specific questions devised by the research group. Table 1 SMART prompt format Clinical scenarios Theoretical questions Patient’s inquiries Seeker I'm a head and neck surgeon with over 15 years of experience, working in a tertiary-level hospital, and I specialize in [type of surgery relevant to the scenario] I am a fifth-year resident specializing in head and neck surgery. I am a patient Mission I need your advice on how to manage a patient I'm treating I need your help to prepare for my final residency exam. I need medical advice for a problem that has been diagnosed in me. AI role You are the world's leading expert in [type of surgery relevant to the scenario] You are a full professor of head and neck surgery at the most prestigious university in the world. The world's leading expert in head and neck surgery. Register correct and specialized scientific language. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information correct and specialized scientific language. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information clear and understandable language even for a non-expert. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information Targeted question [non-contextualized question] [non-contextualized question] [non-contextualized question] On July 31, 2024, the uncontextualized question and the question formatted with the SMART prompt were entered into ChatGPT-4o 14 by the same researcher. For each question, a new browser window was opened in incognito mode to ensure a fresh session. The responses generated by the AI were collected in a Word document, randomized, and then provided to reviewers for evaluation [Supplementary Table 1]. 2.3 Evaluation protocol A pool of three reviewers, all head and neck surgeons with over 15 years of experience, conducted a blind evaluation of the AI-generated responses to both formats of questions. The assessment was performed using the QAMAI instrument, a validated tool designed to evaluate the quality of healthcare information provided by AI chatbots 15 . The QAMAI instrument assesses various aspects of the information quality, including accuracy, clarity, relevance, completeness, source quality and usefulness. Each dimension is scored on a scale from 1 to 5, with the total QAMAI score being the sum of these individual scores, allowing a maximum possible score of 30 for each response. 2.4 Statistical analysis The statistical analysis was conducted by a blinded researcher using Jamovi software version 2.3.18.0, a freeware and open statistical software available online at www.jamovi.org . Descriptive statistics for quantitative variables are reported as median [interquartile range (IQR)] or as mean ± standard deviation. The difference between the QAMAI scores obtained from the responses to the two question formats was evaluated using the Wilcoxon signed-rank test for paired samples. In all cases, the level of statistical significance was set at p < 0.05 with a 95% confidence interval. 3. RESULTS Overall, the questions formulated using the SMART prompt yielded significantly higher quality responses compared to those that were not contextualized (QAMAI score 24 [IQR 21.8–25] versus 27.5 [IQR 25–29]; p < 0.001). The scores were significantly better across all evaluated items: accuracy (4 [IQR 3.75-4] versus 4.5 [IQR 4–5], p < 0.001), clarity (4 [IQR 4–5] versus 5 [IQR 4–5], p = 0.002), relevance (4 [IQR 3–4] versus 4.5 [IQR 4–5], p < 0.001), completeness (4 [IQR 3–4] versus 5 [IQR 4–5]; p < 0.001), and usefulness (4 [IQR 3.75-4] versus 5 [IQR 4–5], p < 0.001). When analyzing the different types of questions, clinical scenarios presented using the SMART prompt achieved significantly higher QAMAI scores compared to non-contextualized questions (QAMAI score 25 [IQR 23.3–25.3] versus 28 [IQR 26–29]; p < 0.001). All individual items reported significantly better scores, except for the quality of sources (4 [IQR 4–5] versus 4 [IQR 4–5], p < 0.256) [Table 2 ]. Table 2 Results of the assessment of the quality of the responses. Non-contestualized prompt Median [IQR] SMART prompt Median [IQR] p-value Clinical scenarios Accuracy 4 [4–5] 5 [4–5] < 0.001 Clarity 4 [4–5] 5 [5–5] 0.013 Relevance 4 [4–4] 5 [4–5] 0.001 Completeness 4 [3–4] 5 [4–5] < 0.001 Sources 4 [4–5] 4 [4–5] 0.256 Usefulness 4 [3.75-4] 5 [4–5] < 0.001 QUAMAI score 25 [23.3–25.3] 28 [26–29] < 0.001 Theoretical questions Accuracy 4 [3-4.25] 4 [4–5] 0.013 Clarity 4 [4–5] 5 [4–5] 0.187 Relevance 4 [3–5] 4 [3.75-5] 0.495 Completeness 4 [3–5] 4 [4–5] 0.181 Sources 2 [2–3] 4 [4–4] < 0.001 Usefulness 4 [3-4.25] 4 [4–5] 0.113 QUAMAI score 23 [17.8–25.3] 25 [23–28] 0.011 Patient’s questions Accuracy 4 [4–4] 5 [4–5] 0.003 Clarity 4.5 [4–5] 5 [4–5] 0.182 Relevance 4 [3.75-4] 4.5 [4–5] 0.002 Completeness 4 [3.75-4] 4.5 [4–5] 0.004 Sources 4 [3–4] 4 [4–5] 0.006 Usefulness 4 [4–4] 5 [4–5] 0.011 QUAMAI score 24 [23–25] 28 [25-29.3] 0.001 Regarding theoretical questions, the responses obtained using the SMART prompt were significantly better (QAMAI score 23 [IQR 17.8–25.3] versus 25 [IQR 23–28]; p < 0.011). In terms of individual items, the questions submitted with the SMART prompt were significantly more accurate (4 [IQR 3-4.25] versus 4 [IQR 4–5]; p = 0.013) and cited significantly more reliable sources (2 [IQR 2–3] versus 4 [IQR 4–4]; p < 0.001). The other items did not show significant differences between the two formats [Table 2 ]. Finally, for patient inquiries, ChatGPT-4o performed significantly better when using the SMART prompt compared to non-contextualized prompts (QAMAI score 24 [IQR 23–25] versus 28 [IQR 25-29.3]; p = 0.001). In the individual items, the AI achieved significantly better scores across all categories when using the SMART prompt, except for clarity (4.5 [IQR 4–5] versus 5 [IQR 4–5], p = 0.182) [Table 2 ]. 4. DISCUSSION The integration of AI in medicine holds the potential for revolutionizing patient care, clinical decision-making, and medical education 17 , 18 . AI systems, like chatbots, can process vast amounts of data quickly, provide tailored information, and enhance accessibility to medical knowledge. However, the use of AI also carries significant risks, including the potential for disseminating inaccurate information, over-reliance on automated systems, and challenges in ensuring patient privacy and data security 19 , 20 . Balancing these potentials with the associated risks is crucial as AI becomes increasingly embedded in healthcare practices 21 . The findings of this study highlight the critical role that prompt construction plays in enhancing the quality of AI chatbot responses, particularly within the context of head and neck surgery. The introduction of the SMART prompt format has demonstrated a significant improvement in the overall performance of AI-driven chatbots, as evidenced by the superior QAMAI scores across its various dimensions. One of the most striking outcomes of this study is the clear benefit of providing detailed and structured contextual information through the SMART prompt format allowing the AI to generate responses that are more precise and tailored to the specific needs of the user. The ability of the AI to deliver more accurate and relevant responses in these complex scenarios underscores the importance of a well-structured prompt that encapsulates the necessary context for effective interaction. In clinical scenarios, the responses generated with the SMART prompt not only delve deeper into the details of the case but also provide a hierarchical order of possible treatments, including specific drug dosages and a more critical analysis of the scenarios. The AI considers all the provided information to offer more precise and appropriate therapeutic recommendations. This detailed approach is crucial in ensuring that the AI can assist healthcare professionals in making well-informed clinical decisions. The study also reveals that the benefits of the SMART prompt are not uniform across all types of questions. While clinical scenarios and patient inquiries showed marked improvements, the difference in performance for theoretical questions, although statistically significant, was less pronounced. This could suggest that the inherent nature of theoretical questions, which may require less contextualization, limits the degree to which prompt formatting can influence AI performance. However, even in these cases, the SMART prompt led to better source quality, indicating that the structured approach encourages the AI to reference more reliable and relevant information, a crucial factor in scientific and educational contexts 22 , 23 . In patient inquiries, the SMART prompt also proved effective, enabling the AI to adjust the context appropriately by providing correct information in an accessible language. This is particularly important as more patients are likely to turn to AI for healthcare information in the future. Ensuring that patients can obtain the most reliable information possible will be essential. As AI continues to play an increasingly prominent role in patient care and medical education, the importance of optimizing AI-human interactions cannot be overstated. The SMART prompt format provides a practical framework for enhancing the efficacy of AI tools, ensuring that they can deliver high-quality, contextually appropriate information. While the study provides robust evidence supporting the utility of the SMART prompt, there are some limitations that should be acknowledged. The study was conducted using a specific AI platform (ChatGPT-4o), and while the results are promising, they may not be fully generalizable to other AI systems. Additionally, the study focused on a relatively narrow field of medicine; future research should explore the applicability of the SMART prompt format across different medical specialties and AI platforms. Further research is also needed to refine the SMART prompt format and explore its potential integration into AI systems as a standard feature. This could involve developing automated tools that assist users in constructing effective prompts, thereby democratizing access to high-quality AI interactions across a broader range of healthcare professionals. 5. CONCLUSIONS In conclusion, the study demonstrates that the quality of AI-generated responses in healthcare can be significantly enhanced through the use of a structured, context-rich prompt format such as SMART. This approach not only improves the accuracy and relevance of the information provided but also contributes to more reliable and effective use of AI in clinical practice. As AI continues to evolve, the development of best practices for prompt formulation will be essential to maximize the potential of these technologies in improving patient care and supporting medical professionals. Declarations ACNOWLEDGEMENTS None ETHICAL APPROVAL Ethical committee approval was not required for this study as it did not involve any patients. FUNDING: none CONFLICT OF INTEREST: none AUTHORS CONTRIBUTIONS: Luigi Angelo Vaira: conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval. Jerome R. Lechien: conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval. Vincenzo Abbate: data collection, data curation, revision of the original and final draft, final approval. Guido Gabriele: data collection, data curation, revision of the original and final draft, final approval. Andrea Frosolini: development of the methodology, data curation, revision of the original and final draft, final approval. Andrea De Vito: development of the methodology, data curation, revision of the original and final draft, final approval. Antonino Maniaci: literature review, data curation, revision of the original and final draft, final approval. Miguel Mayo-Yáñez: literature review, data curation, revision of the original and final draft, final approval. Paolo Boscolo-Rizzo: literature review, data curation, revision of the original and final draft, final approval. Alberto Maria Saibene: development of the methodology, data curation, revision of the original and final draft, final approval. Fabio Maglitto: development of the methodology, data curation, revision of the original and final draft, final approval. Giovanni Salzano: data collection, data curation, revision of the original and final draft, final approval. Gianluigi Califano: statistical analysis, data curation, revision of the original and final draft, final approval. Stefania Troise: development of the methodology, data curation, revision of the original and final draft, final approval. Carlos Miguel Chiesa-Estomba: development of the methodology, data curation, revision of the original and final draft, final approval. Giacomo De Riu: supervision, conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval. References Topol EJ (2019) High-performance medicine: the convergence of human and artificial intelligence. Nat Med 25:44–56 Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA et al (2024) Accuracy of ChatGPT-Generated Information on Head and Neck and Oromaxillofacial Surgery: A Multicenter Collaborative Analysis. Otolaryngol Head Neck Surg 170:1492–1503 Lechien JR, Naunheim MR, Maniaci A, Radulesco T, Saibene AM, Chiesa-Estomba CM, Vaira LA (2024) Performance and Consistency of ChatGPT-4 Versus Otolaryngologists: A Clinical Case Series. Otolaryngol Head Neck Surg 170:1519–1526 Banerjee S, Dunn P, Conard S, Ali A (2024) Mental Health Applications of Generative AI and Large Language Modeling in the United States. Int J Environ Res Public Health 21:910 Chen A, Chen DO, Tian L Benchmarking the symptom-checking capabilities of ChatGPT for a broad range of diseases. J Am Med Inf Assoc 2023 Dec 18:ocad245. 10.1093/jamia/ocad245 . Epub ahead of print. Fraser H, Crossland D, Bacher I, Ranney M, Madsen T, Hilliard R (2023) Comparison of Diagnostic and Triage Accuracy of Ada Health and WebMD Symptom Checkers, ChatGPT, and Physicians for Patients in an Emergency Department: Clinical Data Analysis Study. JMIR Mhealth Uhealth 11:e49995 Saibene AM, Allevi F, Calvo-Henriquez C, Maniaci A, Mayo-Yáñez M, Paderno A et al (2024) Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation. Eur Arch Otorhinolaryngol 281:1835–1841 De Vito A, Geremia N, Marino A, Bavaro DF, Caruana G, Meschiari M et al (2024) Assessing ChatGPT's theoretical knowledge and prescriptive accuracy in bacterial infections: a comparative study with infectious diseases residents and specialists. Infection. 12. 10.1007/s15010-024-02350-6 . Epub ahead of print Anisha SA, Sen A, Bain C (2024) Evaluating the Potential and Pitfalls of AI-Powered Conversational Agents as Humanlike Virtual Health Carers in the Remote Management of Noncommunicable Diseases: Scoping Review. J Med Internet Res 26:e56114 Nadarzynski T, Miles O, Cowie A, Ridge D (2019) Acceptability of artificial intelligence (AI)-led chatbot services in healthcare: A mixed-methods study. Digit Health 5:2055207619871808 Campbell DJ, Estephan LE, Sina EM, Mastrolonardo EV, Alapati R, Amin DR, Cottrill EE (2024) Evaluating ChatGPT Responses on Thyroid Nodules for Patient Education. Thyroid 34(3):371–377 Lee TJ, Campbell DJ, Rao AK, Hossain A, Elkattawy O, Radfar N et al (2024) Evaluating ChatGPT Responses on Atrial Fibrillation for Patient Education. Cureus 16(6):e61680 Raza A, Latif M, Umer Farooq M, Adnan Baig M, Ali Akhtar M, Waseemullah (2023) Enabling Context-based AI in Chatbots for conveying Personalized Interdisciplinary Knowledge to Users. Eng Technol Appl Sci 13:12231–12236 ChatGPT-4o (2023) Available online: https://openai.com/blog/chatgpt Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA et al (2024 May) Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 4. 10.1007/s00405-024-08710-0 Epub ahead of print The jamovi project (2022) Jamovi. (version 2.3) [Computer Software]. Retrieved from https://www.jamovi.org Laranjo L, Dunn AG, Tong HL, Kocaballi AB, Chen J, Bashir R et al (2018) Conversational agents in healthcare: a systematic review. J Am Med Inf Ass 25:1248–1258 Sallam M (2023) ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthc (Basel) 11:887 Dave T, Athaluri SA, Singh S (2023) ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 10.3389/frai.2023.1169595 Cheng K, Li Z, He Y, Guo Q, Lu Y, Gu S, Wu H (2023) Potential use of artificial intelligence in infectious disease: take ChatGPT as an example. Ann Biomed Eng 51:1130–1135 Lee JC, Hamill CS, Shnayder Y, Buczek E, Kakarala K, Bur AM (2024) Exploring the Role of Artificial Intelligence Chatbots in Preoperative Counseling for Head and Neck Cancer Surgery. Laryngoscope 134:2757–2761 Frosolini A, Franz L, Benedetti S, Vaira LA, de Filippis C, Gennaro P et al (2023) Assessing the accuracy of ChatGPT references in head and neck and ENT disciplines. Eur Arch Otorhinolaryngol 280:5129–5133 Lechien JR, Briganti G, Vaira LA (2024) Accuracy of ChatGPT-3.5 and – 4 in providing scientific references in otolaryngology-head and neck surgery. Eur Arch Otorhinolaryngol 281:2159–2165 Additional Declarations The authors declare no competing interests. Supplementary Files Supplementarytable1.docx Supplementary table 1. Questions and answers given by ChatGPT-4. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4953716","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":343373854,"identity":"0cdb691e-0a89-4cc6-9531-9fc5d206f75c","order_by":0,"name":"Luigi Angelo Vaira","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5ElEQVRIiWNgGAWjYBACPgbGBoYEIIONgYHxAZDm4SOkhQ1JC7MBSAsbYS1IbAl0EexaJJKbXzxgOCzPJ334WeXXHDsZNgbmh49u4NWS2GaRwHDYsI0vzey27LZkoMPYjI1zCGgxSGBIY2zjYTC7LbmNGaiFh02aGC32bTzs34olt9UTpaX5QQKDTWIbD48Z48dth4nQwvOwjSHBwCYZqKVYmnHbcR42ZgJ+4WdPf/zxR4WE7fwe9o0ff26rtudnb374GJ8WsNsYDCAsZh4wiV85WMkHGIvxB2HVo2AUjIJRMAIBAIN+OyK0U8YRAAAAAElFTkSuQmCC","orcid":"","institution":"University of Sassari","correspondingAuthor":true,"prefix":"","firstName":"Luigi","middleName":"Angelo","lastName":"Vaira","suffix":""},{"id":343373855,"identity":"dded5495-aa63-4d36-bdfb-34a3ac49b503","order_by":1,"name":"Jerome R. Lechien","email":"","orcid":"","institution":"University of Mons","correspondingAuthor":false,"prefix":"","firstName":"Jerome","middleName":"R.","lastName":"Lechien","suffix":""},{"id":343373856,"identity":"2b6ab47b-3ab2-4db9-ade1-c9dd48bac214","order_by":2,"name":"Vincenzo Abbate","email":"","orcid":"","institution":"University of Naples Federico II","correspondingAuthor":false,"prefix":"","firstName":"Vincenzo","middleName":"","lastName":"Abbate","suffix":""},{"id":343373857,"identity":"6e466292-e53c-4537-803c-e2b660e6376f","order_by":3,"name":"Guido Gabriele","email":"","orcid":"","institution":"University of Siena","correspondingAuthor":false,"prefix":"","firstName":"Guido","middleName":"","lastName":"Gabriele","suffix":""},{"id":343373858,"identity":"fc54aea5-5b07-4ad2-992b-d5d0d23e5926","order_by":4,"name":"Andrea Frosolini","email":"","orcid":"","institution":"University of Siena","correspondingAuthor":false,"prefix":"","firstName":"Andrea","middleName":"","lastName":"Frosolini","suffix":""},{"id":343373859,"identity":"7bca5bcd-ff1b-4104-975e-d4b69b873d95","order_by":5,"name":"Andrea De Vito","email":"","orcid":"","institution":"University of Sassari","correspondingAuthor":false,"prefix":"","firstName":"Andrea","middleName":"","lastName":"De Vito","suffix":""},{"id":343373860,"identity":"3ebcea5e-916b-487e-98c9-63f3908f422a","order_by":6,"name":"Antonino Maniaci","email":"","orcid":"","institution":"University of Enna","correspondingAuthor":false,"prefix":"","firstName":"Antonino","middleName":"","lastName":"Maniaci","suffix":""},{"id":343373861,"identity":"82f26cef-7948-4c4c-97e2-1c95e745d824","order_by":7,"name":"Miguel Mayo Yanez","email":"","orcid":"","institution":"Complexo Hospitalario Universitario A Coruña","correspondingAuthor":false,"prefix":"","firstName":"Miguel","middleName":"Mayo","lastName":"Yanez","suffix":""},{"id":343373862,"identity":"5d3f1714-9cfe-4fcc-a001-0992440fd37d","order_by":8,"name":"Paolo Boscolo-Rizzo","email":"","orcid":"","institution":"University of Trieste","correspondingAuthor":false,"prefix":"","firstName":"Paolo","middleName":"","lastName":"Boscolo-Rizzo","suffix":""},{"id":343373863,"identity":"25363285-afb5-4671-93b7-6edde47dcd2d","order_by":9,"name":"Alberto Maria Saibene","email":"","orcid":"","institution":"University of Milan","correspondingAuthor":false,"prefix":"","firstName":"Alberto","middleName":"Maria","lastName":"Saibene","suffix":""},{"id":343373864,"identity":"8886a128-6a8f-4970-b0c4-4ab2d0f7a939","order_by":10,"name":"Fabio Maglitto","email":"","orcid":"","institution":"University of Bari","correspondingAuthor":false,"prefix":"","firstName":"Fabio","middleName":"","lastName":"Maglitto","suffix":""},{"id":343373865,"identity":"8183875d-8028-4d9b-93dd-c7b0ffee4da4","order_by":11,"name":"Giovanni Salzano","email":"","orcid":"","institution":"University of Bari","correspondingAuthor":false,"prefix":"","firstName":"Giovanni","middleName":"","lastName":"Salzano","suffix":""},{"id":343373866,"identity":"09456936-764e-461f-b755-431209df17d8","order_by":12,"name":"Gianluigi Califano","email":"","orcid":"","institution":"University of Naples Federico II","correspondingAuthor":false,"prefix":"","firstName":"Gianluigi","middleName":"","lastName":"Califano","suffix":""},{"id":343373867,"identity":"ad3ffc62-2fb1-4be3-8dee-9973e2e94b10","order_by":13,"name":"Stefania Troise","email":"","orcid":"","institution":"University of Naples Federico II","correspondingAuthor":false,"prefix":"","firstName":"Stefania","middleName":"","lastName":"Troise","suffix":""},{"id":343373868,"identity":"6a881586-4063-4539-95ae-00032945730e","order_by":14,"name":"Carlos Miguel Chiesa-Estomba","email":"","orcid":"","institution":"Donostia University","correspondingAuthor":false,"prefix":"","firstName":"Carlos","middleName":"Miguel","lastName":"Chiesa-Estomba","suffix":""},{"id":343373869,"identity":"195f9ac6-314a-45e0-8b1b-b4de7280e804","order_by":15,"name":"Giacomo De Riu","email":"","orcid":"","institution":"University of Sassari","correspondingAuthor":false,"prefix":"","firstName":"Giacomo","middleName":"","lastName":"De Riu","suffix":""}],"badges":[],"createdAt":"2024-08-21 19:52:52","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-4953716/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4953716/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":63135426,"identity":"54170605-c6ec-4e3c-9b30-f8f8418e08c1","added_by":"auto","created_at":"2024-08-23 14:15:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":489226,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4953716/v1/8c608880-f01d-47fe-a134-5c53f70b7cfb.pdf"},{"id":63135418,"identity":"8975981a-80bd-45e9-a41c-127e71f35626","added_by":"auto","created_at":"2024-08-23 14:15:25","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":122223,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary table 1. Questions and answers given by ChatGPT-4.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Supplementarytable1.docx","url":"https://assets-eu.researchsquare.com/files/rs-4953716/v1/d1de26d14590a8dc32dd5c43.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eEnhancing AI Chatbot Responses in Healthcare: The SMART Prompt Structure in Head and Neck Surgery\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. INTRODUCTION","content":"\u003cp\u003eIn recent years, the integration of Artificial Intelligence (AI) within the healthcare sector has rapidly expanded, offering unprecedented opportunities to enhance patient care, streamline administrative processes, and support clinical decision-making\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Among the various AI applications, conversational agents, commonly known as chatbots, have gained significant attention for their potential to interact with users in a natural language format. These AI-driven systems are being increasingly utilized for a wide range of purposes, including symptom checking, patient triage, mental health support, and health information dissemination\u003csup\u003e\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe effectiveness of AI chatbots in healthcare, however, largely depends on the quality of the interactions they facilitate\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. A key factor influencing these interactions is the way users formulate their prompts or queries\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Unlike traditional search engines, where users might rely on keyword-based queries, interactions with AI chatbots often involve more complex, context-rich language, which can significantly affect the responses generated by these systems. This is particularly crucial in healthcare settings, where the accuracy, clarity, and relevance of information can have profound implications on patient outcomes. Providing adequate context to an AI chatbot is crucial for obtaining accurate and contextually appropriate responses. The quality of AI-generated responses is significantly influenced by the amount and relevance of the information provided in the prompt\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. When a prompt lacks sufficient context, the AI is forced to make broader assumptions, which can lead to generalized or less accurate responses. Conversely, when a prompt includes detailed and specific information, such as patient history, symptoms, or the specific surgical context, the AI can tailor its responses more precisely to the needs of the user.\u003c/p\u003e \u003cp\u003eDespite the growing use of AI chatbots, there is limited number of research exploring how the structuring of prompts influences the quality of responses in healthcare-related contexts\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Understanding this relationship is essential for optimizing the design and deployment of AI systems in healthcare, ensuring they provide reliable and useful information to users. Moreover, as AI continues to evolve, developing best practices for prompt formulation could enhance the overall user experience and effectiveness of AI-driven healthcare services.\u003c/p\u003e \u003cp\u003eThis study aims to investigate the impact of prompt construction on the quality of AI chatbot responses specifically within the context of head and neck surgery.\u003c/p\u003e"},{"header":"2. MATERIALS AND METHODS","content":"\u003cp\u003eThis study was conducted as part of an international collaborative research project involving young researchers from the International Federation of Otorhinolaryngology Societies and the Italian Society of Maxillofacial Surgery. The consortium was established in February 2023, with the purpose of exploring the potential applications, assessing the reliability, and identifying possible risks associated with artificial intelligence (AI) platforms in the field of head and neck surgery. Sixteen researchers from 11 European centers participated in this study. The requirement for an ethical review and approval was waived because the study did not include any analysis of humans or animals.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Prompt development\u003c/h2\u003e \u003cp\u003eFor the specific purpose of developing the prompt to be used in this study, a multidisciplinary team was constituted. This team included two head and neck surgeons, a linguist, and a computer engineer. Their objective was to create a prompt that would effectively guide the AI in providing accurate and contextually appropriate responses. The prompt format was developed through multiple rounds of testing and refinement. Initial versions of the prompt were tested using various clinical scenarios provided by the head and neck surgeons. Feedback was collected and analyzed to identify any areas where the AI responses were suboptimal. Adjustments were made to the prompt format to improve clarity, specificity, and relevance. The final prompt structure was validated by applying it across different clinical cases to ensure its effectiveness in generating high-quality AI responses.\u003c/p\u003e \u003cp\u003eThis process resulted in the creation of a prompt format identified by the acronym SMART, which stands for Seeker, Mission, AI Role, Register, and Targeted Question.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSeeker\u003c/b\u003e: Refers to the identity and perspective of the user inquiring (i.e. head and neck surgeon, general practitioner, medical student, patient). The prompt must communicate the seeker\u0026rsquo;s role and the clinical context to ensure the AI understands the expertise level and specific needs of the user.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eMission\u003c/b\u003e: Defines the purpose of the inquiry. This involves clearly stating the objective or problem that the AI is expected to address. The mission statement ensures the AI focuses on the most relevant information.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eAI Role\u003c/b\u003e: Describes the expected role of the AI in the interaction, which may range from providing information to offering clinical advice. Clearly defining the AI\u0026rsquo;s role helps set expectations for the type and depth of the response.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eRegister\u003c/b\u003e: Involves the tone and style of language to be used in the AI\u0026rsquo;s response. The prompt should guide the AI to use language that is appropriate for the target audience, ensuring clarity and avoiding ambiguity.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTargeted Question\u003c/b\u003e: Refers to the specific question or query posed to the AI. This question should be direct and focused, designed to elicit a detailed and accurate response.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Study design\u003c/h2\u003e \u003cp\u003eA separate group of three head and neck surgeons, distinct from those involved in the prompt development, was tasked with creating a set of 24 questions. These questions were divided into three groups of eight questions each, covering different types of inquiries: clinical scenarios, theoretical questions, and patient inquiries. The questions encompassed three main areas of head and neck surgery: oncology, sinus and nasal surgery, and trauma. The questions were carefully reviewed to eliminate any potential ambiguities, ensuring clarity and precision.\u003c/p\u003e \u003cp\u003eThe questions were formulated in a question format suitable for direct input into the chatbot. Additionally, all questions required the chatbot to provide bibliographic references for the information it utilized in generating responses. The same research group also prepared the contextual information to be included in the SMART prompt [Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e]. The \"Targeted Question\" section of the SMART prompt was then populated with the specific questions devised by the research group.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSMART prompt format\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClinical scenarios\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTheoretical questions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePatient\u0026rsquo;s inquiries\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSeeker\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eI'm a head and neck surgeon with over 15 years of experience, working in a tertiary-level hospital, and I specialize in [type of surgery relevant to the scenario]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI am a fifth-year resident specializing in head and neck surgery.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eI am a patient\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMission\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eI need your advice on how to manage a patient I'm treating\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI need your help to prepare for my final residency exam.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eI need medical advice for a problem that has been diagnosed in me.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAI role\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYou are the world's leading expert in [type of surgery relevant to the scenario]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYou are a full professor of head and neck surgery at the most prestigious university in the world.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eThe world's leading expert in head and neck surgery.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eRegister\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ecorrect and specialized scientific language. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ecorrect and specialized scientific language. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eclear and understandable language even for a non-expert. The information must be based on the most recent and solid scientific evidence. I would like you to also provide me with the bibliographical references from which you draw your information\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTargeted question\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e[non-contextualized question]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e[non-contextualized question]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e[non-contextualized question]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOn July 31, 2024, the uncontextualized question and the question formatted with the SMART prompt were entered into ChatGPT-4o\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e by the same researcher. For each question, a new browser window was opened in incognito mode to ensure a fresh session. The responses generated by the AI were collected in a Word document, randomized, and then provided to reviewers for evaluation [Supplementary Table\u0026nbsp;1].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Evaluation protocol\u003c/h2\u003e \u003cp\u003eA pool of three reviewers, all head and neck surgeons with over 15 years of experience, conducted a blind evaluation of the AI-generated responses to both formats of questions. The assessment was performed using the QAMAI instrument, a validated tool designed to evaluate the quality of healthcare information provided by AI chatbots\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. The QAMAI instrument assesses various aspects of the information quality, including accuracy, clarity, relevance, completeness, source quality and usefulness. Each dimension is scored on a scale from 1 to 5, with the total QAMAI score being the sum of these individual scores, allowing a maximum possible score of 30 for each response.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Statistical analysis\u003c/h2\u003e \u003cp\u003eThe statistical analysis was conducted by a blinded researcher using Jamovi software version 2.3.18.0, a freeware and open statistical software available online at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"http://www.jamovi.org\" target=\"_blank\"\u003ewww.jamovi.org\u003c/a\u003e\u003c/span\u003e\u003cspan address=\"http://www.jamovi.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Descriptive statistics for quantitative variables are reported as median [interquartile range (IQR)] or as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation. The difference between the QAMAI scores obtained from the responses to the two question formats was evaluated using the Wilcoxon signed-rank test for paired samples. In all cases, the level of statistical significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 with a 95% confidence interval.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. RESULTS","content":"\u003cp\u003eOverall, the questions formulated using the SMART prompt yielded significantly higher quality responses compared to those that were not contextualized (QAMAI score 24 [IQR 21.8\u0026ndash;25] versus 27.5 [IQR 25\u0026ndash;29]; \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001). The scores were significantly better across all evaluated items: accuracy (4 [IQR 3.75-4] versus 4.5 [IQR 4\u0026ndash;5], \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001), clarity (4 [IQR 4\u0026ndash;5] versus 5 [IQR 4\u0026ndash;5], \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.002), relevance (4 [IQR 3\u0026ndash;4] versus 4.5 [IQR 4\u0026ndash;5], \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001), completeness (4 [IQR 3\u0026ndash;4] versus 5 [IQR 4\u0026ndash;5]; \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001), and usefulness (4 [IQR 3.75-4] versus 5 [IQR 4\u0026ndash;5], \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003eWhen analyzing the different types of questions, clinical scenarios presented using the SMART prompt achieved significantly higher QAMAI scores compared to non-contextualized questions (QAMAI score 25 [IQR 23.3\u0026ndash;25.3] versus 28 [IQR 26\u0026ndash;29]; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). All individual items reported significantly better scores, except for the quality of sources (4 [IQR 4\u0026ndash;5] versus 4 [IQR 4\u0026ndash;5], p\u0026thinsp;\u0026lt;\u0026thinsp;0.256) [Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of the assessment of the quality of the responses.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNon-contestualized prompt\u003c/p\u003e \u003cp\u003eMedian [IQR]\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSMART prompt\u003c/p\u003e \u003cp\u003eMedian [IQR]\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ep-value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eClinical scenarios\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClarity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [5\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.013\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRelevance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCompleteness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSources\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.256\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUsefulness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3.75-4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQUAMAI score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25 [23.3\u0026ndash;25.3]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e28 [26\u0026ndash;29]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTheoretical questions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3-4.25]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.013\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClarity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.187\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRelevance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [3.75-5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.495\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCompleteness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.181\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSources\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2 [2\u0026ndash;3]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUsefulness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3-4.25]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.113\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQUAMAI score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e23 [17.8\u0026ndash;25.3]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e25 [23\u0026ndash;28]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.011\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePatient\u0026rsquo;s questions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.003\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClarity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4.5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.182\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRelevance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3.75-4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4.5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCompleteness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3.75-4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4.5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.004\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSources\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [3\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.006\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUsefulness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4 [4\u0026ndash;4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5 [4\u0026ndash;5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.011\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQUAMAI score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e24 [23\u0026ndash;25]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e28 [25-29.3]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eRegarding theoretical questions, the responses obtained using the SMART prompt were significantly better (QAMAI score 23 [IQR 17.8\u0026ndash;25.3] versus 25 [IQR 23\u0026ndash;28]; p\u0026thinsp;\u0026lt;\u0026thinsp;0.011). In terms of individual items, the questions submitted with the SMART prompt were significantly more accurate (4 [IQR 3-4.25] versus 4 [IQR 4\u0026ndash;5]; p\u0026thinsp;=\u0026thinsp;0.013) and cited significantly more reliable sources (2 [IQR 2\u0026ndash;3] versus 4 [IQR 4\u0026ndash;4]; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). The other items did not show significant differences between the two formats [Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFinally, for patient inquiries, ChatGPT-4o performed significantly better when using the SMART prompt compared to non-contextualized prompts (QAMAI score 24 [IQR 23\u0026ndash;25] versus 28 [IQR 25-29.3]; p\u0026thinsp;=\u0026thinsp;0.001). In the individual items, the AI achieved significantly better scores across all categories when using the SMART prompt, except for clarity (4.5 [IQR 4\u0026ndash;5] versus 5 [IQR 4\u0026ndash;5], p\u0026thinsp;=\u0026thinsp;0.182) [Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e"},{"header":"4. DISCUSSION","content":"\u003cp\u003eThe integration of AI in medicine holds the potential for revolutionizing patient care, clinical decision-making, and medical education\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. AI systems, like chatbots, can process vast amounts of data quickly, provide tailored information, and enhance accessibility to medical knowledge. However, the use of AI also carries significant risks, including the potential for disseminating inaccurate information, over-reliance on automated systems, and challenges in ensuring patient privacy and data security\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Balancing these potentials with the associated risks is crucial as AI becomes increasingly embedded in healthcare practices\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe findings of this study highlight the critical role that prompt construction plays in enhancing the quality of AI chatbot responses, particularly within the context of head and neck surgery. The introduction of the SMART prompt format has demonstrated a significant improvement in the overall performance of AI-driven chatbots, as evidenced by the superior QAMAI scores across its various dimensions. One of the most striking outcomes of this study is the clear benefit of providing detailed and structured contextual information through the SMART prompt format allowing the AI to generate responses that are more precise and tailored to the specific needs of the user. The ability of the AI to deliver more accurate and relevant responses in these complex scenarios underscores the importance of a well-structured prompt that encapsulates the necessary context for effective interaction.\u003c/p\u003e \u003cp\u003eIn clinical scenarios, the responses generated with the SMART prompt not only delve deeper into the details of the case but also provide a hierarchical order of possible treatments, including specific drug dosages and a more critical analysis of the scenarios. The AI considers all the provided information to offer more precise and appropriate therapeutic recommendations. This detailed approach is crucial in ensuring that the AI can assist healthcare professionals in making well-informed clinical decisions.\u003c/p\u003e \u003cp\u003eThe study also reveals that the benefits of the SMART prompt are not uniform across all types of questions. While clinical scenarios and patient inquiries showed marked improvements, the difference in performance for theoretical questions, although statistically significant, was less pronounced. This could suggest that the inherent nature of theoretical questions, which may require less contextualization, limits the degree to which prompt formatting can influence AI performance. However, even in these cases, the SMART prompt led to better source quality, indicating that the structured approach encourages the AI to reference more reliable and relevant information, a crucial factor in scientific and educational contexts\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn patient inquiries, the SMART prompt also proved effective, enabling the AI to adjust the context appropriately by providing correct information in an accessible language. This is particularly important as more patients are likely to turn to AI for healthcare information in the future. Ensuring that patients can obtain the most reliable information possible will be essential. As AI continues to play an increasingly prominent role in patient care and medical education, the importance of optimizing AI-human interactions cannot be overstated. The SMART prompt format provides a practical framework for enhancing the efficacy of AI tools, ensuring that they can deliver high-quality, contextually appropriate information.\u003c/p\u003e \u003cp\u003eWhile the study provides robust evidence supporting the utility of the SMART prompt, there are some limitations that should be acknowledged. The study was conducted using a specific AI platform (ChatGPT-4o), and while the results are promising, they may not be fully generalizable to other AI systems. Additionally, the study focused on a relatively narrow field of medicine; future research should explore the applicability of the SMART prompt format across different medical specialties and AI platforms.\u003c/p\u003e \u003cp\u003eFurther research is also needed to refine the SMART prompt format and explore its potential integration into AI systems as a standard feature. This could involve developing automated tools that assist users in constructing effective prompts, thereby democratizing access to high-quality AI interactions across a broader range of healthcare professionals.\u003c/p\u003e"},{"header":"5. CONCLUSIONS","content":"\u003cp\u003eIn conclusion, the study demonstrates that the quality of AI-generated responses in healthcare can be significantly enhanced through the use of a structured, context-rich prompt format such as SMART. This approach not only improves the accuracy and relevance of the information provided but also contributes to more reliable and effective use of AI in clinical practice. As AI continues to evolve, the development of best practices for prompt formulation will be essential to maximize the potential of these technologies in improving patient care and supporting medical professionals.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eACNOWLEDGEMENTS\u003c/p\u003e\n\u003cp\u003eNone\u003c/p\u003e\n\u003cp\u003eETHICAL APPROVAL\u003c/p\u003e\n\u003cp\u003eEthical committee approval was not required for this study as it did not involve any patients.\u003c/p\u003e\n\u003cp\u003eFUNDING: none\u003c/p\u003e\n\u003cp\u003eCONFLICT OF INTEREST: none\u003c/p\u003e\n\u003cp\u003eAUTHORS CONTRIBUTIONS:\u003c/p\u003e\n\u003cp\u003eLuigi Angelo Vaira: conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval.\u003c/p\u003e\n\u003cp\u003eJerome R. Lechien:\u0026nbsp;conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval.\u003c/p\u003e\n\u003cp\u003eVincenzo Abbate: data collection, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eGuido Gabriele: data collection, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eAndrea Frosolini: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eAndrea De Vito: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eAntonino Maniaci: literature review, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eMiguel Mayo-Y\u0026aacute;\u0026ntilde;ez: literature review, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003ePaolo Boscolo-Rizzo: literature review, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eAlberto Maria Saibene: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eFabio Maglitto: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eGiovanni Salzano: data collection, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eGianluigi Califano: statistical analysis, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eStefania Troise: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eCarlos Miguel Chiesa-Estomba: development of the methodology, data curation, revision of the original and final draft, final approval.\u003c/p\u003e\n\u003cp\u003eGiacomo De Riu: supervision, conceptualization of the work, development of the methodology, data curation, writing the original draft, writing the final draft, final approval.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTopol EJ (2019) High-performance medicine: the convergence of human and artificial intelligence. Nat Med 25:44\u0026ndash;56\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA et al (2024) Accuracy of ChatGPT-Generated Information on Head and Neck and Oromaxillofacial Surgery: A Multicenter Collaborative Analysis. Otolaryngol Head Neck Surg 170:1492\u0026ndash;1503\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLechien JR, Naunheim MR, Maniaci A, Radulesco T, Saibene AM, Chiesa-Estomba CM, Vaira LA (2024) Performance and Consistency of ChatGPT-4 Versus Otolaryngologists: A Clinical Case Series. Otolaryngol Head Neck Surg 170:1519\u0026ndash;1526\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBanerjee S, Dunn P, Conard S, Ali A (2024) Mental Health Applications of Generative AI and Large Language Modeling in the United States. Int J Environ Res Public Health 21:910\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen A, Chen DO, Tian L Benchmarking the symptom-checking capabilities of ChatGPT for a broad range of diseases. J Am Med Inf Assoc 2023 Dec 18:ocad245. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/jamia/ocad245\u003c/span\u003e\u003cspan address=\"10.1093/jamia/ocad245\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub ahead of print.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFraser H, Crossland D, Bacher I, Ranney M, Madsen T, Hilliard R (2023) Comparison of Diagnostic and Triage Accuracy of Ada Health and WebMD Symptom Checkers, ChatGPT, and Physicians for Patients in an Emergency Department: Clinical Data Analysis Study. JMIR Mhealth Uhealth 11:e49995\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaibene AM, Allevi F, Calvo-Henriquez C, Maniaci A, Mayo-Y\u0026aacute;\u0026ntilde;ez M, Paderno A et al (2024) Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation. Eur Arch Otorhinolaryngol 281:1835\u0026ndash;1841\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Vito A, Geremia N, Marino A, Bavaro DF, Caruana G, Meschiari M et al (2024) Assessing ChatGPT's theoretical knowledge and prescriptive accuracy in bacterial infections: a comparative study with infectious diseases residents and specialists. Infection. 12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s15010-024-02350-6\u003c/span\u003e\u003cspan address=\"10.1007/s15010-024-02350-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Epub ahead of print\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnisha SA, Sen A, Bain C (2024) Evaluating the Potential and Pitfalls of AI-Powered Conversational Agents as Humanlike Virtual Health Carers in the Remote Management of Noncommunicable Diseases: Scoping Review. J Med Internet Res 26:e56114\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNadarzynski T, Miles O, Cowie A, Ridge D (2019) Acceptability of artificial intelligence (AI)-led chatbot services in healthcare: A mixed-methods study. Digit Health 5:2055207619871808\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCampbell DJ, Estephan LE, Sina EM, Mastrolonardo EV, Alapati R, Amin DR, Cottrill EE (2024) Evaluating ChatGPT Responses on Thyroid Nodules for Patient Education. Thyroid 34(3):371\u0026ndash;377\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee TJ, Campbell DJ, Rao AK, Hossain A, Elkattawy O, Radfar N et al (2024) Evaluating ChatGPT Responses on Atrial Fibrillation for Patient Education. Cureus 16(6):e61680\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRaza A, Latif M, Umer Farooq M, Adnan Baig M, Ali Akhtar M, Waseemullah (2023) Enabling Context-based AI in Chatbots for conveying Personalized Interdisciplinary Knowledge to Users. Eng Technol Appl Sci 13:12231\u0026ndash;12236\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChatGPT-4o (2023) Available online: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://openai.com/blog/chatgpt\u003c/span\u003e\u003cspan address=\"https://openai.com/blog/chatgpt\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA et al (2024 May) Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 4. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00405-024-08710-0\u003c/span\u003e\u003cspan address=\"10.1007/s00405-024-08710-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003eEpub ahead of print\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThe jamovi project (2022) Jamovi. (version 2.3) [Computer Software]. Retrieved from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.jamovi.org\u003c/span\u003e\u003cspan address=\"https://www.jamovi.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaranjo L, Dunn AG, Tong HL, Kocaballi AB, Chen J, Bashir R et al (2018) Conversational agents in healthcare: a systematic review. J Am Med Inf Ass 25:1248\u0026ndash;1258\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSallam M (2023) ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthc (Basel) 11:887\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDave T, Athaluri SA, Singh S (2023) ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/frai.2023.1169595\u003c/span\u003e\u003cspan address=\"10.3389/frai.2023.1169595\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng K, Li Z, He Y, Guo Q, Lu Y, Gu S, Wu H (2023) Potential use of artificial intelligence in infectious disease: take ChatGPT as an example. Ann Biomed Eng 51:1130\u0026ndash;1135\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee JC, Hamill CS, Shnayder Y, Buczek E, Kakarala K, Bur AM (2024) Exploring the Role of Artificial Intelligence Chatbots in Preoperative Counseling for Head and Neck Cancer Surgery. Laryngoscope 134:2757\u0026ndash;2761\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrosolini A, Franz L, Benedetti S, Vaira LA, de Filippis C, Gennaro P et al (2023) Assessing the accuracy of ChatGPT references in head and neck and ENT disciplines. Eur Arch Otorhinolaryngol 280:5129\u0026ndash;5133\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLechien JR, Briganti G, Vaira LA (2024) Accuracy of ChatGPT-3.5 and \u0026ndash;\u0026thinsp;4 in providing scientific references in otolaryngology-head and neck surgery. Eur Arch Otorhinolaryngol 281:2159\u0026ndash;2165\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of Sassari","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"ChatGPT, artificial intelligence, AI, prompt engineering, maxillofacial surgery, otorhinolaryngology","lastPublishedDoi":"10.21203/rs.3.rs-4953716/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4953716/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjective.\u003c/h2\u003e \u003cp\u003eTo evaluate the impact of prompt construction on the quality of AI chatbot responses in the context of head and neck surgery.\u003c/p\u003e\u003ch2\u003eStudy design.\u003c/h2\u003e \u003cp\u003eObservational and evaluative study.\u003c/p\u003e\u003ch2\u003eSetting.\u003c/h2\u003e \u003cp\u003eInternational collaboration involving 16 researchers from 11 European centers specializing in head and neck surgery.\u003c/p\u003e\u003ch2\u003eMethods.\u003c/h2\u003e \u003cp\u003eA total of 24 questions, divided into clinical scenarios, theoretical questions, and patient inquiries, were developed. These questions were inputted into ChatGPT-4o both with and without the use of a structured prompt format, known as SMART (Seeker, Mission, AI Role, Register, Targeted Question). The AI-generated responses were evaluated by experienced head and neck surgeons using the QAMAI instrument, which assesses accuracy, clarity, relevance, completeness, source quality, and usefulness.\u003c/p\u003e\u003ch2\u003eResults.\u003c/h2\u003e \u003cp\u003eThe responses generated using the SMART prompt scored significantly higher across all QAMAI dimensions compared to those without contextualized prompts. Median QAMAI scores for SMART prompts were 27.5 (IQR 25\u0026ndash;29) versus 24 (IQR 21.8\u0026ndash;25) for unstructured prompts (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Clinical scenarios and patient inquiries showed the most significant improvements, while theoretical questions also benefited but to a lesser extent. The AI's source quality improved notably with the SMART prompt, particularly in theoretical questions.\u003c/p\u003e\u003ch2\u003eConclusions.\u003c/h2\u003e \u003cp\u003eThe study suggests that the structured SMART prompt format significantly enhances the quality of AI chatbot responses in head and neck surgery. This approach improves the accuracy, relevance, and completeness of AI-generated information, underscoring the importance of well-constructed prompts in clinical applications. Further research is warranted to explore the applicability of SMART prompts across different medical specialties and AI platforms.\u003c/p\u003e","manuscriptTitle":"Enhancing AI Chatbot Responses in Healthcare: The SMART Prompt Structure in Head and Neck Surgery","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-23 14:15:20","doi":"10.21203/rs.3.rs-4953716/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9c0fb87a-dcf2-4523-a536-9748d9956938","owner":[],"postedDate":"August 23rd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":36384337,"name":"Artificial Intelligence and Machine Learning"},{"id":36384338,"name":"Otorhinolaryngology"}],"tags":[],"updatedAt":"2024-08-23T14:15:20+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-23 14:15:20","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4953716","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4953716","identity":"rs-4953716","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.