A Comparative Analysis of Five AI Chatbots in Providing Patient Education on Smile Design | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Comparative Analysis of Five AI Chatbots in Providing Patient Education on Smile Design Bahadır Ezmek, Hasan Alper Uyar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8210813/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: This study aimed to evaluate and compare the accuracy, quality, readability, understandability, and actionability of responses provided by five AI chatbots—Microsoft Copilot, ChatGPT-4, ChatGPT-5, Google Gemini, and Claude Sonet 4.5—to patient questions about smile design and anterior aesthetic dental procedures. Method: Twenty-eight patient-oriented questions were collected from Reddit and Quora. A volunteer asked these questions to the five AI chatbots on the same day in a blinded order. Each response was recorded and coded to maintain anonymity. Two prosthodontists independently assessed the responses for accuracy using a 5-point Likert scale, quality using the Global Quality Scale (GQS), and understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT-P). Readability was measured with Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability was calculated using Cohen’s kappa. Statistical analyses were performed using Kruskal-Wallis tests for non-parametric data and ANOVA for normally distributed readability scores, with p < 0.05 considered statistically significant. Results: Significant differences were observed in accuracy (p = 0.013) and quality (p < 0.001) among the chatbots. ChatGPT-5 had lower accuracy than Google Gemini (p = 0.017) and Claude Sonet 4.5 (p = 0.041) and lower quality than all other chatbots (p < 0.001). Readability differed significantly (FRE: p = 0.004; FKGL: p < 0.001), with ChatGPT-5 responses requiring the highest reading level. PEMAT-P scores also showed significant differences in understandability and actionability (p < 0.001), with ChatGPT-5 displaying lower scores than other chatbots. Microsoft Copilot, ChatGPT-4, and Google Gemini generally provided higher-quality, more understandable, and actionable information, while ChatGPT-5 and Claude Sonet 4.5 showed limitations. Most chatbot responses were above an eighth-grade reading level, which may challenge general patient comprehension. Conclusion: AI chatbots vary considerably in the quality and usefulness of information they provide for complex dental procedures like smile design. While some models deliver accurate and comprehensible responses, others may produce lower-quality, less actionable content. Despite high understandability in most responses, high reading levels and low actionability could limit patient comprehension and effective decision-making. Care should be taken when patients rely on AI chatbots for dental education, and further improvements are needed to enhance reliability, readability, and actionable guidance. AI chatbots smile design accuracy assessment health information quality readability understandability and actionability Figures Figure 1 Background Today, a beautiful smile has evolved beyond being merely an aesthetic feature to become an important element of communication that directly affects a person's social interactions, self-confidence, and psychosocial well-being. Individuals' perceptions of their own dental aesthetics have a significant impact on their social behavior and dental self-confidence ( 1 ). The increased visibility of faces in the digital age and the frequent presentation of “perfect” smiles on social media contribute to individuals evaluating their own appearance more often and turning to cosmetic procedures ( 2 , 3 ). It has been reported that an aesthetically pleasing and healthy smile has a powerful impact on first impressions and social perception, particularly in young adults, which is related to the quality of life and social participation ( 4 , 5 ). For this reason, today, smile design is not merely a clinical application but has gained importance as a comprehensive intervention to improve the individual's psychosocial well-being and social functioning. Although digital smile design applications improve information sharing and joint decision-making by showing patients their possible aesthetic results before treatment, they also strengthen patient–dentist communication through tools such as facial-tooth ratio analysis and mock-up previews. ( 6 , 7 ) However, many patients still feel the need to seek a second clinical opinion for a number of reasons. There are many reasons for this trend. Perhaps the most important of these are the irreversible nature of the treatment, high aesthetic expectations, and high treatment costs ( 8 ). Although simulations obtained with digital smile design contribute to high patient satisfaction, patients prefer to discuss alternative treatment plans and risk-benefit analyses with other clinicians before fully approving the treatment. Therefore, despite the transparency and visualization advantages provided by digital smile design, seeking a second opinion during the treatment process remains a common and justified practice in cosmetic dentistry. Today patients not only seek second opinions from different clinicians, but also heavily rely on online forums and artificial intelligence (AI) chatbots when making treatment decisions. Especially in fields such as cosmetic dentistry, individuals are examining patient experiences shared on health forums and engaging in dialogue with large language model (LLM)-based chatbots like ChatGPT to learn more about smile design and potential complications. Cross-sectional studies show that users utilize LLM-based chatbots to obtain health information and benefit from these tools at a rate of approximately 21% ( 9 ). However, clinicians express concerns that these tools may affect the trust dynamics between patients and clinicians ( 10 ). For this reason, in cosmetic dentistry, face-to-face consultations between the patient and the dentist and digital simulation processes play a crucial role in shaping patients' expectations and perceptions, thereby establishing trust between the patient and the dentist. AI chatbots are developed using LLM and are based on the transformer architecture. This structure evaluates the context between words in a sentence in a multi-layered manner through the “self-attention” mechanism. First, the model learns to predict the next word during pre-training with massive text data. Then, during the fine-tuning phase, it is optimized for dialogue or question-answering using human-labeled data. Many models are guided to provide safer, more user-friendly responses through “Reinforcement Learning from Human Feedback”. When a user asks a question, the model predicts the most appropriate response based on context and statistical probabilities ( 11 , 12 ). Some users may find AI chatbots appealing due to the informative and accessible content they provide. However, debates continue regarding the reliability and accuracy of AI chatbots. AI chatbots may sometimes exhibit a tendency to “hallucinate” (generate unrealistic or incorrect information) ( 13 ). Some studies highlight the potential of chatbots to provide inaccurate information when generating medical content ( 14 – 17 ). Previous studies in several specialties of dentistry have also determined that AI chatbots can provide users with incorrect information ( 17 – 21 ). In addition, it has been reported that the “Reinforcement Learning from Human Feedback” process does not completely eliminate the risk of the model being manipulated or misdirected ( 22 ). The potential to generate inaccurate, outdated, or biased content, particularly in health information, is a critical concern for patient safety ( 13 ). Whether patients obtain information from forums or AI chatbots, the readability of the text is as important as its accuracy in enabling patients to draw the necessary conclusions. The readability of a text is found by its Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) ( 19 , 21 ). Readability tools score the text by assessing the complexity of vocabulary and syntax. However, they do not directly measure understandability and actionability ( 21 ). The Patient Education Materials Assessment Tool (PEMAT-P) was designed to assess the understandability of a patient information/education text. PEMAT-P also has an additional seven questions, which are used to evaluate actionability ( 21 ). Research on the accuracy, quality, comprehensibility, and applicability of AI chatbots in answering dentistry-related questions has become an increasingly common and actively explored topic. ( 18 – 21 , 23 – 28 ). However, the varying results reported across these studies make it difficult to generalize that any AI chatbot can consistently provide high-quality responses across all topics. Therefore, it remains uncertain whether AI chatbots can reliably serve as comprehensive tools for patient education. Additionally, to the best of our knowledge, there are no studies evaluating the performance of AI chatbots in smile design. Thus, this study aimed to assess and compare the accuracy and quality of the information provided by five AI chatbots to patient questions. Additionally, the readability, understandability, and actionability of responses of chatbots were also evaluated. The null hypotheses of this study were that ( 1 ) accuracy and quality of responses would not differ among five different AI chatbots nor existing literature, and ( 2 ) readability, understandability, and actionability of responses would not differ. Methods In this study, the methodology of a previous study was used to determine patient concerns and questions regarding smile design and the method ( 21 ). Searches were conducted on the popular question-and-answer sites Reddit ( https://www.reddit.com/r/asksdentists/ ) and Quora ( https://www.quora.com/ ) using the search keywords “smile design” and “anterior esthetic dental restorations”. Questions were recorded. Using these records, a total of 28 questions were prepared under the headings of general questions, treatment procedure, care and maintenance, and functional and aesthetic concerns ( 18 ). The questions were prepared using as little dental terminology as possible, so that they appear to be asked by the patient rather than a dental professional (Table 1 ). Table 1 Defined questions about smile design and anterior dental esthetics General Questions 1. Which restoration type is best for my front tooth: direct laminate, indirect laminate, or full crown? 2. What happens if I have short teeth or a “gummy smile”? 3. Which materials give the most natural result — metal-ceramic, full ceramic, or composite? 4. Are these cosmetic gum and tooth treatments permanent or reversible? 5. How much will this kind of smile design cost, and how long will it last? 6. Is gum health important before doing veneers or crowns? Treatment & Procedure 7. What exactly happens during a gum contouring? 8. How long does it take for gums to heal after reshaping? 9. What if the gum surgery is done incorrectly — can it be fixed later? 10. Will gum surgery change the way my teeth look immediately? 11. For laminates and full crowns, how much tooth needs to be removed? 12. How many appointments are needed for veneers and gum reshaping? 13. How do dentists decide where my new gum line should be? 14. If my gums grow back or shrink after the surgery, will it affect my restorations? 15. Can I combine gum reshaping and laminate placement in the same visit? Care & Maintenance 16. How can I prevent my gums from becoming inflamed again after cosmetic treatment? 17. Do veneers or crowns require special flossing techniques around the gumline? 18. How often should I have professional cleaning or maintenance after smile design? 19. Can gums recede after veneers or crowns, and what can I do to prevent that? 20. Can whitening or mouthwash irritate gums after cosmetic procedures? Functional And Aesthetic Concerns 21. How much do gums affect the beauty of my smile? 22. Will reshaping my gums make my teeth look longer and more even? 23. What if the gum levels become uneven — can it ruin the symmetry of my smile? 24. Can poorly shaped gums make veneers or crowns look fake? 25. Will gum surgery change how my lips move when I smile? 26. Can gum color be adjusted or treated if it looks uneven after surgery? 27. How long will it take for my gums to look completely natural again? 28. Does gum surgery have any risk of affecting tooth sensitivity or root exposure? In this study, since no human subjects or personal data were used, ethical approval was not needed in accordance with the Declaration of Helsinki. The research involved only the evaluation of AI chatbot responses to pre-defined questions, ensuring that no identifiable personal information was collected or analyzed. A volunteer who did not participate in the study asked the questions to the Microsoft Copilot, Google Gemini, ChatGPT-4, ChatGPT-5, and Claude Sonet 4.5 chatbots in the same order and on the same day on November 10, 2025. To prevent the answer from affecting earlier questions, a new chat window was opened for each question. All responses were saved in a Microsoft Word file with the original formatting preserved. Each chatbot’s responses document was number-coded before being handled to authors for blinded assessment. For blinded assessment, the volunteer also prepared a document for the evaluation of each question (See Supplementary file 1). Two prosthodontists (authors) rated all responses’ accuracy, quality, understandability, and actionability separately. For accuracy assessment, a five-point Likert scale ranging from 1 (the chatbot’s answer is completely incorrect), to 2 (the chatbot’s answer contains more incorrect items than correct items), to 3 (the chatbot’s answer contains an equal balance of correct and incorrect items), to 4 (the chatbot’s answer contains more correct items than incorrect items), and to 5 (the chatbot’s answer is completely correct) was used ( 18 ). The Global Quality Scale (GQS) with a 5-point scoring system was used to evaluate the quality of the responses ( 18 , 21 ). FRE and FKGL scores were automatically calculated using a freely available online tool ( https://goodcalculators.com/flesch-kincaid-calculator/ ). This online tool calculates FRE and FKGL using the following formulas: FRE score = 206.835 − 1.015 (total number of words/total number of sentences) − 84.6 (total number of syllables/total number of words). FKGL score = 0.39 (total number of words/total number of sentences) + 11.8 (total number of syllables/total number of words) − 15.59. FRE scores can be interpreted by referencing established conversion tables to estimate the reader’s required education level, whereas FKGL directly reports the minimum grade level needed to comprehend the text ( 18 , 19 ). PEMAT-P was used to assess the understandability and actionability of responses of AI chatbots. The PEMAT-P consists of 17 questions for understandability and 7 questions for actionability. First, prosthodontists read the PEMAT-P manual and examined examples related to scoring. Each author independently rated all responses as 0 for “disagree”, 1 for “agree”, and N/A for “not applicable” ( 18 , 21 ). Since the responses did not have visual aids, the relevant questions of PEMAT-P were marked “not applicable” and excluded from the evaluation (See Supplementary file 2). Statistical analysis Data were statistically analyzed using a statistical software (SPSS, version 22.0, IBM Corp., Armonk, NY). Cohen’s kappa was calculated to evaluate the inter-rater agreement between the two researchers. Inter-rater agreement for the 5-point Likert scale, GQS, PEMAT-P/Understandability, and PEMAT-P/Actionability was assessed using the Cohen’s kappa statistic. Cohen’s kappa values of 0.00-0.20 were considered slight agreement, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81-1.00 almost perfect. Data normality was checked using the Kolmogorov-Smirnov test, and Levene’s test was used to assess the homogeneity of the variances. 5-point Likert scale, GQS, PEMAT-P/Understandability, and PEMAT-P/Actionability were non-normally distributed. These variables were analyzed using the Kruskal-Wallis test. Differences among the AI chatbot responses were determined using Dunn’s test. FRE and FKGL scores, normally distributed, were analyzed using a one-way analysis of variance (ANOVA) and Tukey’s HSD post hoc tests. Statistical significance was set at p < 0.05. Results Cohen’s kappa values indicated consistent inter-rater reliability across all measures. The CQS showed moderate agreement (κ = 0.55), while the PEMAT-P understandability demonstrated substantial agreement (κ = 0.73). Both the 5-point Likert scale and PEMAT-P actionability assessments showed almost perfect agreement (κ = 0.85 and κ = 0.84, respectively). The results are shown in Table 2 . Table 2 Comparison of Accuracy, Quality, FRE, FKGL, PEMAT-P/Understandability, and PEMAT-P/Actionability scores according to AI chatbots. Microsoft Copilot ChatGPT-4 Google Gemini ChatGPT-5 Claude Sonet 4.5 p Accuracy Five-point Likert scale Median (min-max) 5 ( 3 – 5 ) ab 5 ( 4 – 5 ) ab 5 ( 4 – 5 ) a 5 ( 3 – 5 ) b 5 ( 4 – 5 ) a 0. 013 Quality GQS scores Median (min-max) 5 ( 3 – 5 ) ac 5 ( 4 – 5 ) a 5 ( 4 – 5 ) a 4 ( 3 – 5 ) b 3 ( 3 – 5 ) c < 0.001 FRE Mean (Standard deviation) 43.26 (10.6) ab 41.27 (10.83) ab 49.05 (6.59) b 37.33 (11.95) a 42.26 (14.25) ab 0.004 FKGL Mean (Standard deviation) 9.02 (1.55) ab 9.27 (1.44) ab 9.76 (1.03) b 12.78 (1.98) c 8.52 (1.97) a < 0.001 PEMAT-P Understandability Median (min-max) 84.61 (71.42–100) a 84.61 (57.14–92.85) ab 83.33 (58.33–92.85) ab 75 (45.45- 90) b 83.33 (42.85–92.3) b < 0.001 PEMAT-P Actionability Median (min-max) 40 (0–80) a 20 (0–80) a 40(0–60) a 20 (0–60) b 40(0–80) a < 0.001 (GQS: Global Quality Scale, FRE: Flesch Reading Ease, FKGL: Flesch-Kincaid Grade Level, PEMAT-P: Patient Education Materials Assessment Tool) A statistically significant difference was found in the median accuracy scores among AI chatbots, X 2 (4, 280) = 12.658, p = 0. 013. The accuracy scores of ChatGPT-5 were lower than both Google Gemini and Claude Sonet 4.5 (p = 0.017 and p = 0.041, respectively). A statistically significant difference was also found in the median quality scores among AI chatbots, X 2 (4, 280) = 104.525, p = 0.001. The quality scores of ChatGPT-5 were significantly different from all other AI chatbots (p < 0.001). There were also significant differences between the quality scores of Claude Sonet 4.5 and both ChatGPT-4 (p = 0.042) and Google Gemini (p = 0.025). One-way ANOVA revealed significant differences in FRE scores of AI chatbot responses (F(4,135) = 4.047, p = 0.004). Tukey’s HSD Test for multiple comparisons found that the mean FRE scores were significantly different between Google Gemini and ChatGPT-5 (p = 0.001). There was no statistically significant difference in mean FRE scores between other AI chatbots (p > 0.05). One-way ANOVA also revealed significant differences in FKGL scores of AI chatbot responses (F(4,135) = 30.838, p < 0.001). Higher FKGL scores were calculated in ChatGPT-5 than other AI chatbots (p 0.05). As listed in Table 3 , the majority of AI chatbot responses were written at an 8th-grade level or higher. All ChatGPT-5 responses corresponded to an education level of 8th grade or higher. The proportions of responses below an 8th-grade level were 25% for Microsoft Copilot, 17.85% for ChatGPT-4, 3.6% for Google Gemini, and 42.85% for Claude Sonet 4.5. Table 3 Number and percentage of the Flesch-Kincaid Grade Level (FKGL) of the AI chatbot responses FKGL (Readability Level) Microsoft Copilot n (%) ChatGPT-4 n (%) Google Gemini n (%) ChatGPT-5 n (%) Claude Sonet 4.5 n (%) 5 (Very Easy) 0 (0) 0 (0) 0 (0) 0 (0) 0 (0) 6 (Easy) 3 (10.7) 1 (3.6) 0 (0) 0 (0) 8 (28.6) 7 (Fairly Easy) 4 (14.3) 4 (14.3) 1 (3.6) 0 (0) 4 (14.3) 8–9 (Standard) 13 (46.4) 14 (50) 16 (57.1) 3 (10.7) 10 (35.7) 10–12 (Fairly Difficult) 7 ( 25 ) 9 (32.1) 11 (39.3) 12 (42.9) 6 (21.4) 13–16 (Difficult) 1 (3.6) 0 (0) 0 (0) 11 (39.3) 0 (0) College Graduate (Very Difficult) 0 (0) 0 (0) 0 (0) 2 (7.1) 0 (0) Figure 1 shows the distribution of PEMAT-P understandability and actionability scores. Kruskal–Wallis analysis showed that PEMAT-P understandability scores of AI chatbots were significantly different, X 2 (4, 280) = 21.726, p < 0.001. The median PEMAT-P/Understandability score of Microsoft Copilot was significantly different than ChatGPT-5 (p < 0.001) and Claude Sonet 4.5 (p = 0.016). Kruskal–Wallis analysis also revealed that the PEMAT-P Actionability scores of AI chatbots were significantly different, X 2 (4, 280) = 46.727, p < 0.001. Dunn’s pairwise comparisons revealed that the PEMAT-P Actionability scores of ChatGPT-5 were significantly different from the PEMAT-P Actionability scores of other AI chatbots. Discussion This study compared the performance of five different AI chatbots using multiple evaluation tools, including the 5-point Likert scale, the GQS, PEMAT-P Understandability, PEMAT-P Actionability, and readability indices (FRE and FKGL). The results suggested that chatbot-generated health information shows considerable variability. Overall, Microsoft Copilot, ChatGPT-4, and Google Gemini tended to provide higher-quality and more comprehensible responses, while ChatGPT-5 and Claude Sonet 4.5 showed lower performance in accuracy and quality. Therefore, the first null hypothesis was rejected. The readability, understandability, and actionability of AI chatbot responses were also significantly different. The second null hypothesis was also rejected. In multi-step smile design procedures that include surgical interventions, the accuracy and quality of patient information help ensure that patients correctly understand the treatment process, which in turn supports improved treatment outcomes. Whether the information being accurate or of poor quality can directly influence the patient’s understanding, expectations, and overall satisfaction with the treatment. Therefore, the responses provided by AI chatbots, which patients may easily use to obtain information about the procedure, should also be accurate and of high quality. However, in this study, the significant differences observed in the 5-point Likert scale and GQS ratings reflect variations in the overall reliability, completeness, and usefulness of the AI chatbot responses. Higher-scoring models appeared to provide more structured explanations, fewer irrelevant details, and better alignment with established clinical knowledge. The high accuracy and quality of the Microsoft Copilot, ChatGPT-4, and Google Gemini observed in this study are consistent with those reported in a previous study ( 29 ). In contrast, lower-scoring AI chatbots (ChatGPT-5 and Claude Sonet 4.5) often demonstrated oversimplification, occasional inconsistencies, and limited contextual depth. When patients are not accurately informed, they may develop unfounded fears or concerns about the safety of smile design procedures, which may cause them to hesitate or refuse to undergo the necessary procedures. Additionally, conflicting information between AI chatbots and dentists can undermine trust and prevent patients from following professional advice or seeking timely treatment for any future dental procedures ( 19 ). Some studies have reported that the scientific accuracy of AI chatbot responses increases depending on the sophistication of the Natural Language Processing (NLP) system and the use of larger datasets ( 19 , 30 ). Despite their advanced NLP architectures and large training datasets, several reasons may explain why ChatGPT-5 and Claude Sonet 4.5 show reduced accuracy and quality. The responses generated by ChatGPT-5 were excessively brief and lacked the level of informational depth expected even for patient-oriented content. This response pattern may have resulted from several factors, including safety filters and alignment mechanisms. The safety filters and alignment mechanisms of ChatGPT-5 can make their answers more conservative and less specific, which lowers perceived quality in complex or technical domains ( 13 ). On the other hand, although Claude Sonet 4.5 gave long answers, some of them were scientifically incorrect. This is likely because the model can produce fluent but wrong information (hallucinations), which may seem convincing even when it is not accurate ( 15 , 16 ). Consistent with the results of this study, Tuzlalı et al.( 17 ) also mentioned the limited accuracy and completeness of Claude responses for questions about dental implants. This shows that there is a risk of errors or missing information in some responses. Improving the suitability of an AI model for dentistry topics can be achieved by providing additional training to the base model. This process creates a dentistry-focused expert model and may significantly increase its accuracy ( 20 ). An accurate text with high readability can help patients understand medical information more quickly, participate more actively in decision-making, and follow treatment procedures more accurately. Patient education materials should be written at or below the 8th grade-level ( 19 ). Materials written above this level may reduce comprehension and limit effective patient engagement. In this study, within these criteria, the lowest readability scores were observed for ChatGPT-5 (with an FRE of 37.33 and an FKGL of 12.78). The high FKGL scores of ChatGPT‑5 responses may be attributed to the use of technical terminology and complex sentence structures, which increase reading difficulty ( 31 ). The five AI chatbots evaluated in this study produced responses with mean FRE scores between 37.33 and 49.05 and mean FKGL scores between 8.52 and 12.78. When evaluating the FRE scores, no statistically significant differences were observed among the AI chatbots except for ChatGPT‑5, whereas the lowest FKGL scores were found in Microsoft Copilot, ChatGPT‑4, and Claude Sonet 4.5. In the literature on dental procedures, the readability scores of AI chatbot responses vary across studies. Özcivelek and Özcan ( 18 ), in their study on maxillofacial prosthesis, reported that the FKGL scores of the four AI chatbots (DeepSeek-R1, Chat GPT-o1, ChatGPT-4, and Dental GPT) varied between 9.42 and 10.70, corresponding to a ‘fairly difficult’ reading level. Özdemir et al. ( 32 ), in their study on oral habits, reported average FKGL scores of 10.47 for ChatGPT‑4, 9.53 for Google Gemini, and 9.61 for Microsoft Copilot. ChatGPT‑4 responses were found to be more complex and less readable than those of Gemini and Copilot. Esmailpour et al.( 20 ) also reported that ChatGPT-4 (10.45) produced higher FKGL scores compared to Microsoft Copilot (8.38) and Google Gemini (7.82).In contrast to Özdemir et al. and Esmailpour et al, our study found no significant differences in FKGL scores between ChatGPT‑4, Microsoft Copilot, and Google Gemini. Güven et al.( 21 ), in their study on traumatic dental injuries, also found no significant difference between ChatGPT-4 and Google Gemini, although both produced responses at a college-level readability. In a study evaluating responses on endodontic iatrogenic events of four AI chatbots (ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4), the authors reported that the generated answers were classified as ‘fairly difficult’ (high school) to ‘difficult’ (college) to read ( 33 ). As seen in previous studies, when the same AI chatbots are asked questions on different dental topics, the readability of their responses can vary considerably. It is due to questions about specific procedures or technical topics making the AI chatbot use more complex words or longer sentences, which raise the reading level, while general patient-focused questions produce simpler and easier-to-read text ( 18 , 32 ). In addition, adding prompts that include simple language rules (short sentences, common words, active sentence structure) to the questions can improve readability ( 34 ). However, individuals who do not routinely use AI chatbots and who are not aware of their limitations may not always be able to apply such prompts to improve the readability of the information provided. Readability alone may not ensure that patients understand dental procedures or can make informed decisions. Therefore, assessing the responses using PEMAT‑P is needed to evaluate the understandability and actionability. In the present study, ChatGPT-5 showed the lowest PEMAT-P understandability scores (75%), which was consistent with its poorer readability results. The median of PEMAT-P understandability scores of the other AI chatbots ranged from 83.33% to 84.61%. The literature recommends a 70% threshold for both PEMAT-P understandability and actionability ( 18 , 35 , 36 ). The PEMAT-P user guide also states that materials scoring above this level can be interpreted as “more understandable and more actionable. In previous studies ( 18 , 21 , 35 , 37 – 39 ), most AI chatbots show PEMAT-P understandability scores of 70% or higher for patient-education responses related to dental topics. In this study, consistent with the literature, most of the understandability scores of AI chatbot responses were found to be above the 70% threshold (Fig. 1 ). However, most actionability scores were found to be below this threshold. Additionally, ChatGPT-5 was found to have lower actionability scores compared to the other AI chatbots (Table 2 ). Several studies ( 18 , 21 , 37 – 39 ) in the literature support our finding that AI chatbot responses show high understandability but low actionability. The reason for lower actionability of AI-generated patient materials may be due to the use of medical terminology and general, non-guiding statements (for example, step-by-step action plans, checklists, and clear recommendations) aimed at preventing errors ( 40 ). In this study, the accuracy and quality of information provided to patients by AI chatbots in complex dental procedures involving multiple procedures, such as smile design, were evaluated. As in other studies, significant differences were found between the AI chatbots in this study. In this study, the identification of the lowest accuracy and quality values in ChatGPT-5 and Claude Sonnet 4.5 reveals that some models still have significant limitations in complex dental decision-making processes. This study has also revealed that factors such as the readability and actionability of AI chatbot responses may limit their use in patient education. However, there are several limitations that may affect the results of this study. First, the study was designed so that a new chat window was used for each question. Each question is entered separately without allowing follow-up questions or explanations, which naturally limits the depth and realism of chatbot interactions. This approach did not assess chatbots' ability to improve, contextualize, and elaborate responses through multi-turn dialogues. Additionally, preparing questions in a structured template format (diagnosis → explanation → options → follow-up) can also improve the response performance of AI chatbots. However, this method was not preferred by the authors, considering the limited intellectual potential of an average individual to use the chatbot in this sequence. Additionally, for questions related to clinical scenarios, criteria such as accuracy and quality may be influenced by the individual judgment of the authors, which can affect the assessment results. In our study, 5-point Likert and Global Quality Scales were used, and high inter-rater reliability results were found on the Likert scale (κ = 0.85) and moderate results on the GQS (κ = 0.55). This situation shows that a certain degree of variability is inevitable in human evaluations. Hybrid evaluation models that combine expert assessments with AI-powered content scoring models can increase consistency and reduce bias ( 17 ). Furthermore, the study evaluated only five publicly available chatbots as of November 2025. The performance and capabilities of AI chatbots are rapidly improving with model updates and fine-tuning iterations. This situation is one of the reasons why studies conducted at different time intervals yield different results. Another limitation of the study is that it was conducted only in English. AI chatbot responses in different languages may differ in terms of performance. AI applications have great potential to develop more accessible and readable online content for patient education. In addition, they can be a step towards increasing patient satisfaction through education. Future studies may evaluate artificial intelligence technologies specialized in dental topics that can create visual aids to answer patient questions. Conclusion This study assessed the accuracy, quality, readability, understandability, and actionability of responses from five AI chatbots for dental questions, including complex procedures like smile design. ChatGPT-5 and Claude Sonet 4.5 had the lowest accuracy and quality scores, which show limits in handling multi-step dental cases. Most responses were understandable, but their readability often needed a high reading level, which may make them harder for general patients to follow. Actionability scores were generally low. Abbreviations AI: Artificial intelligence LLM: Large Language Model FRE: Flesch Reading Ease FKGL: Flesch-Kincaid Grade Level PEMAT-P: Patient Education Materials Assessment Tool GQS: Global Quality Scale ANOVA: One-way analysis of variance NLP: Natural Language Processing Declarations Ethics approval and consent to participate This study did not require ethical approval in accordance with the Declaration of Helsinki, as it involved only evaluation of AI chatbot responses and did not include human subjects or personal data. Consent for publication Not applicable. Data availability The datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare no competing interests. Funding Not applicable. Authors' contributions Conceptualization: B.E. and H.A.U.; Methodology: B.E. and H.A.U.; Formal analysis and investigation B.E. and H.A.U.; Writing—original draft preparation: B.E. and H.A.U.; Writing—review and editing: B.E.; Funding acquisition: B.E and H.A.U.; Laboratory resources: - Supervision: B.E. Acknowledgements The authors would like to thank Dr. Berna Kavuncu Ezmek for her contributions in obtaining and editing the AI chatbot responses. References Afroz S, Rathi S, Rajput G, Rahman SA. Dental esthetics and its impact on psycho-social well-being and dental self confidence: A campus based survey of north indian university students. J Indian Prosthodont Soc. 2013;13(4):455–60. Abbasi MS, Lal A, Das G, Salman F, Akram A, Ahmed AR, et al. Impact of Social Media on Aesthetic Dentistry: General Practitioners’ Perspectives. Healthc. 2022;10(10):1–10. Baik KM, Anbar G, Alshaikh A, Banjar A. Effect of Social Media on Patient’s Perception of Dental Aesthetics in Saudi Arabia. Int J Dent. 2022;2022:4794497. Stojilković M, Gušić I, Berić J, Prodanović D, Pecikozić N, Veljović T, et al. Evaluating the influence of dental aesthetics on psychosocial well-being and self-esteem among students of the University of Novi Sad, Serbia: a cross-sectional study. BMC Oral Health. 2024;24(1):1–11. Armalaite J, Jarutiene M, Vasiliauskas A, Sidlauskas A, Svalkauskiene V, Sidlauskas M, et al. Smile aesthetics as perceived by dental students: A cross-sectional study. BMC Oral Health. 2018;18(1):1–7. Mariam A, Prakash VS, Graduate P. Digital Smile Design: Revolutionizing Aesthetic Dentistry Through Technological Innovation. IJSDR. 2024;9(6):1187–92. Doğan AN. Dijital Gülüş Tasarimi: KullanilanSi̇stemler Ve Avantajlari. Sağlık Bilim Derg. 2020;29(2):138–43. Aktan E. Diş Hekimliğinde Dijital Gülüş Tasarımı Uygulamaları Digital Smile Design Applications in Dentistry. ADO J Clin Sci. 2023;12(3):474–9. Yun HS, Bickmore T. Online Health Information-Seeking in the Era of Large Language Models: Cross-Sectional Web-Based Survey Study. J Med Internet Res. 2025;27. Zeng L, Li Q, Zuo Y, Zhang Y, Li Z. Perceptions and Attitudes of Chinese Oncologists Toward Endorsing AI-Driven Chatbots for Health Information Seeking Among Patients with Cancer: Phenomenological Qualitative Study. J Med Internet Res. 2025;27:1–11. Bridgelall R. Unraveling the mysteries of AI chatbots. Artif Intell Rev. 2024;57:89. González Barman K, Lohse S, de Regt HW. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values. Whose Perspectives? Philos Technol. 2025;38:35. Chow JCL, Li K. Large Language Models in Medical Chatbots: Opportunities, Challenges, and the Need to Address AI Risks. Inf. 2025;16(7):1–24. Lawson McLean A, Hristidis V. Evidence-Based Analysis of AI Chatbots in Oncology Patient Education: Implications for Trust, Perceived Realness, and Misinformation Management. J Cancer Educ. 2025;40(4):482–9. Peykani P, Ramezanlou F, Tanasescu C. Large Language Models: A Structured Taxonomy and Review of Challenges, Limitations, Solutions, and Future Directions. Appl Sci. 2025;15:8103. Zhang W, Zhang J. Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics. 2025;13:856. Tuzlalı M, Baki N, Aral K, Aral CA. Evaluating the performance of AI chatbots in responding to dental implant FAQs: A comparative study. BMC Oral Health. 2025;25:1548. Özcivelek T, Özcan B. Comparative evaluation of responses from DeepSeek-R1, ChatGPT-o1, ChatGPT-4, and dental GPT chatbots to patient inquiries about dental and maxillofacial prostheses. BMC Oral Health. 2025;25:871. Helvacioglu-Yigit D, Demirturk H, Ali K, Tamimi D, Koenig L, Almashraqi A. Evaluating artificial intelligence chatbots for patient education in oral and maxillofacial radiology. Oral Surg Oral Med Oral Pathol Oral Radiol. 2025;139(6):750–9. Esmailpour H, Rasaie V, Babaee Hemmati Y, Falahchai M. Performance of artificial intelligence chatbots in responding to the frequently asked questions of patients regarding dental prostheses. BMC Oral Health. 2025;25:574. Guven Y, Ozdemir OT, Kavan MY. Performance of Artificial Intelligence Chatbots in Responding to Patient Queries Related to Traumatic Dental Injuries: A Comparative Study. Dent Traumatol. 2025;41(3):338–47. Dahlgren Lindström A, Methnani L, Krause L, Ericson P, de Rituerto de Troya ÍM, Coelho Mollo D, et al. Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. Ethics Inf Technol. 2025;27(2):1–13. Mohammad-Rahimi H, Ourang SA, Pourhoseingholi MA, Dianat O, Dummer PMH, Nosrat A. Validity and reliability of artificial intelligence chatbots as public sources of information on endodontics. Int Endod J. 2024;57(3):305–14. Freire Y, Santamaría Laorden A, Orejas Pérez J, Gómez Sánchez M, Díaz-Flores García V, Suárez A. ChatGPT performance in prosthodontics: Assessment of accuracy and repeatability in answer generation. J Prosthet Dent. 2024;131(4):659.e1-659.e6. Babayiğit O, Tastan Eroglu Z, Ozkan Sen D, Ucan Yarkac F. Potential Use of ChatGPT for Patient Information in Periodontology: A Descriptive Pilot Study. Cureus. 2023;15(11):e48518. Daraqel B, Wafaie K, Mohammed H, Cao L, Mheissen S, Liu Y et al. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop [Internet]. 2024;165(6):652–62. Available from: https://linkinghub.elsevier.com/retrieve/pii/S0889540624000593 Aguiar de Sousa R, Costa SM, Almeida Figueiredo PH, Camargos CR, Ribeiro BC, Alves e Silva MRM. Is ChatGPT a reliable source of scientific information regarding third-molar surgery? J Am Dent Assoc [Internet]. 2024;155(3):227–232.e6. Available from: https://linkinghub.elsevier.com/retrieve/pii/S0002817723006815 Kılınç DD, Mansız D. Examination of the reliability and readability of Chatbot Generative Pretrained Transformer’s (ChatGPT) responses to questions about orthodontics and the evolution of these responses in an updated version. Am J Orthod Dentofac Orthop [Internet]. 2024;165(5):546–55. Available from: https://linkinghub.elsevier.com/retrieve/pii/S0889540624000076 Yagci F, Eraslan R, Albayrak H, İpekten F. Accuracy and Reliability of Artificial Intelligence Chatbots as Public Information Sources in Implant Dentistry. Int J Oral Maxillofac Implants. 2025;1–23. Azadi A, Gorjinejad,Fatemeh Mohammad-Rahimi H, Tabrizi R, Alam M, Golkar M. Evaluation of AI-generated responses by different artificial intelligence chatbots to the clinical decision-making case-based questions in oral and maxillofacial surgery. Oral Surg Oral Med Oral Pathol Oral Radiol. 2024;137(6):587–93. Asfuroğlu ZM, Yağar H, Gümüşoğlu E. High accuracy but limited readability of large language model-generated responses to frequently asked questions about Kienböck’s disease. BMC Musculoskelet Disord. 2024;25:879. Özdemir ÖT, Kavan MY, Güven Y. Evaluation of the readability, quality, and accuracy of AI chatbot responses to questions about deleterious oral habits. BMC Oral Health. 2025;25:1812. Taşyürek M, Adıgüzel Ö, Ortaç H. Comparative Evaluation of Responses from ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4 Chatbots to Questions About Endodontic Iatrogenic Events. Healthc. 2025;13:2615. Okuhara T, Furukawa E, Okada H, Yokota R, Kiuchi T. Readability of written information for patients across 30 years: A systematic review of systematic reviews. Patient Educ Couns [Internet]. 2025;135(October 2024):108656. Available from: https://doi.org/10.1016/j.pec.2025.108656 Sivaramakrishnan G, Almuqahwi M, Ansari S, Lubbad M, Alagamawy E, Sridharan K. Assessing the power of AI: a comparative evaluation of large language models in generating patient education materials in dentistry. BDJ Open. 2025;11(1):1–6. Shoemaker SJ, Wolf MSBC. Developement of the Patient Education Material Assessment Tool (PEMAT). Physiol Behav. 2018;176(5):139–48. Spuur K, Currie G, Al-Mousa D, Pape R. Suitability of ChatGPT as a Source of Patient Information for Screening Mammography. Health Promot Pract. 2025;26(4):746–62. Alnsour MM, Alenezi R, Barakat M, AL-Omiri MK. Assessing ChatGPT’s suitability in responding to the public’s inquires on the effects of smoking on oral health. BMC Oral Health. 2025;25(1). Halboub E, Hakami RAM, Khalufi KNA, Hakami SAH, Alhajj MN. Assessment of quality, understandability, actionability, and readability of responses of selected chatbots to the top searched queries about oral cancer. Digit Dent J. 2025;1:100008. Sönmezoğlu Hİ, Güner Sönmezoğlu B, Temel MH, Çakir B. Comprehensibility and readability of selected artificial intelligence chatbots in providing uveitis-related information. Med (Baltim). 2025;104(43):e45135. Additional Declarations No competing interests reported. Supplementary Files Supplemantaryfile1.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8210813","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":551667651,"identity":"b17949f1-f16d-4185-a430-b610549c8cf4","order_by":0,"name":"Bahadır Ezmek","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/0lEQVRIiWNgGAWjYBACAwYeBgYeAwiHsaGCIQHEkCBByxm4FgMCWqAcxsY2IrSYs589+OFNwT0Gfv7TiQ9nzqvLMzjAfPA2D8OffFxaLHvykiXnGBQzSDac3Wy4cdvhYoMDbMnWPAwGlg24HHYgx0CaxyCBweBg7zbJh9sOJG44wGMmDdSC02UG598Y/wZpsT/Mu/3nwzl1QC383/BruZFjBrGFjXcb48YGZpAtbHi1WM54Y2Y5xyCBR+IM72bJGccOF0seZjMGihjj1GLOn2N8482fBDn+/rMbP/bU1OXxHW9+eONNhRzuUIYCHgSTGexgQhpGwSgYBaNgFOADALGxUpXCNDrjAAAAAElFTkSuQmCC","orcid":"","institution":"Health Sciences University","correspondingAuthor":true,"prefix":"","firstName":"Bahadır","middleName":"","lastName":"Ezmek","suffix":""},{"id":551667652,"identity":"4c8d2011-08fd-4945-826c-8b838fb520e8","order_by":1,"name":"Hasan Alper Uyar","email":"","orcid":"","institution":"Health Sciences University","correspondingAuthor":false,"prefix":"","firstName":"Hasan","middleName":"Alper","lastName":"Uyar","suffix":""}],"badges":[],"createdAt":"2025-11-26 08:58:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8210813/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8210813/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":97131701,"identity":"2be8e86d-037c-4cc0-8519-8f46e233fe21","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"jpg","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36855,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/d9f9b0317a17d624bac97d87.jpg"},{"id":97142865,"identity":"22cebefb-62f1-4b9b-853b-f456ce500996","added_by":"auto","created_at":"2025-12-01 10:08:01","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":86368,"visible":true,"origin":"","legend":"","description":"","filename":"AComparativeAnalysisofFiveAIChatbotsinProvidingPatientEducationonSmileDesign.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/e2d509abd7104686f6568d75.docx"},{"id":97131710,"identity":"5f0856ff-ce98-467d-b1ec-24710a1b1d2f","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":17022,"visible":true,"origin":"","legend":"","description":"","filename":"Table1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/1b9444864b54c7c9ad8318e5.docx"},{"id":97131706,"identity":"914a28d0-f0cd-43be-89b4-6fcc2916d193","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16341,"visible":true,"origin":"","legend":"","description":"","filename":"Table2.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/879ce3bc6eef1f0c676103f4.docx"},{"id":97142432,"identity":"64bcda3b-c72b-4054-91c1-be69e36d2291","added_by":"auto","created_at":"2025-12-01 10:07:37","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":15685,"visible":true,"origin":"","legend":"","description":"","filename":"Table3.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/15877579b3c66c1a2e23178a.docx"},{"id":97131702,"identity":"a648d54b-9612-457d-8d58-0ca8c00bd41c","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"json","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5611,"visible":true,"origin":"","legend":"","description":"","filename":"585324a299514b7e9f3edb522d6eae34.json","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/11b2b8361849caa8c1136b8c.json"},{"id":97131709,"identity":"0db11780-9556-4c67-b2f1-b36e14f260c3","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"docx","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28411,"visible":true,"origin":"","legend":"","description":"","filename":"Supplemantaryfile1d1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/fdfb95a0a869b37bb32fef96.docx"},{"id":97142866,"identity":"796cc6c1-65f9-474e-8d71-490965241c25","added_by":"auto","created_at":"2025-12-01 10:08:01","extension":"xml","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":112292,"visible":true,"origin":"","legend":"","description":"","filename":"585324a299514b7e9f3edb522d6eae341enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/b3f8887994b50271065f9238.xml"},{"id":97142969,"identity":"72d01c3d-71fa-4b99-8ea9-73c4ac8df218","added_by":"auto","created_at":"2025-12-01 10:08:10","extension":"jpg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36855,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/e8507a7a1245d193315edc1b.jpg"},{"id":97142423,"identity":"f05c0f78-7293-4901-a5a8-d883dd8cdbfd","added_by":"auto","created_at":"2025-12-01 10:07:37","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":13333,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/24d9bd6337ce65f836c75007.png"},{"id":97131714,"identity":"c9eabba3-acb4-4fad-8d8a-6f285108df95","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"xml","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":111075,"visible":true,"origin":"","legend":"","description":"","filename":"585324a299514b7e9f3edb522d6eae341structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/a834e478b07026d7fa7a1469.xml"},{"id":97131712,"identity":"6d4fd911-e9ff-4977-a4f3-d63ba55ef1dc","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"html","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":117008,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/98c65359b42033f6169f8356.html"},{"id":97131700,"identity":"77cc0886-6db9-4155-b2e7-07c2a111969a","added_by":"auto","created_at":"2025-12-01 08:43:35","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":50078,"visible":true,"origin":"","legend":"\u003cp\u003eBox plots of the PEMAT-P Understandability and Actionability of AI chatbots. The box represents the range from the first to third quartile. The thick horizontal lines inside the boxes are the median values. The reference line shows the 70% threshold for understandability and actionability. Scores at or above this value are considered more understandable and more actionable.\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/0eb4d8454c59c018b2a5eaa5.jpg"},{"id":97145073,"identity":"f979cfd0-f6ef-443e-bc47-ebf34e0edd1d","added_by":"auto","created_at":"2025-12-01 10:12:59","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":740451,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/346570cb-4612-4927-b43d-7cd3cc014869.pdf"},{"id":97142454,"identity":"456f38b0-6424-4cd1-8e5a-c949c2e46103","added_by":"auto","created_at":"2025-12-01 10:07:37","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":28411,"visible":true,"origin":"","legend":"","description":"","filename":"Supplemantaryfile1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8210813/v1/7daee342f667d930888e8d79.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Comparative Analysis of Five AI Chatbots in Providing Patient Education on Smile Design","fulltext":[{"header":"Background","content":"\u003cp\u003eToday, a beautiful smile has evolved beyond being merely an aesthetic feature to become an important element of communication that directly affects a person's social interactions, self-confidence, and psychosocial well-being. Individuals' perceptions of their own dental aesthetics have a significant impact on their social behavior and dental self-confidence (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). The increased visibility of faces in the digital age and the frequent presentation of \u0026ldquo;perfect\u0026rdquo; smiles on social media contribute to individuals evaluating their own appearance more often and turning to cosmetic procedures (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). It has been reported that an aesthetically pleasing and healthy smile has a powerful impact on first impressions and social perception, particularly in young adults, which is related to the quality of life and social participation (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). For this reason, today, smile design is not merely a clinical application but has gained importance as a comprehensive intervention to improve the individual's psychosocial well-being and social functioning.\u003c/p\u003e\u003cp\u003eAlthough digital smile design applications improve information sharing and joint decision-making by showing patients their possible aesthetic results before treatment, they also strengthen patient\u0026ndash;dentist communication through tools such as facial-tooth ratio analysis and mock-up previews. (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e) However, many patients still feel the need to seek a second clinical opinion for a number of reasons. There are many reasons for this trend. Perhaps the most important of these are the irreversible nature of the treatment, high aesthetic expectations, and high treatment costs (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). Although simulations obtained with digital smile design contribute to high patient satisfaction, patients prefer to discuss alternative treatment plans and risk-benefit analyses with other clinicians before fully approving the treatment. Therefore, despite the transparency and visualization advantages provided by digital smile design, seeking a second opinion during the treatment process remains a common and justified practice in cosmetic dentistry.\u003c/p\u003e\u003cp\u003eToday patients not only seek second opinions from different clinicians, but also heavily rely on online forums and artificial intelligence (AI) chatbots when making treatment decisions. Especially in fields such as cosmetic dentistry, individuals are examining patient experiences shared on health forums and engaging in dialogue with large language model (LLM)-based chatbots like ChatGPT to learn more about smile design and potential complications. Cross-sectional studies show that users utilize LLM-based chatbots to obtain health information and benefit from these tools at a rate of approximately 21% (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). However, clinicians express concerns that these tools may affect the trust dynamics between patients and clinicians (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). For this reason, in cosmetic dentistry, face-to-face consultations between the patient and the dentist and digital simulation processes play a crucial role in shaping patients' expectations and perceptions, thereby establishing trust between the patient and the dentist.\u003c/p\u003e\u003cp\u003eAI chatbots are developed using LLM and are based on the transformer architecture. This structure evaluates the context between words in a sentence in a multi-layered manner through the \u0026ldquo;self-attention\u0026rdquo; mechanism. First, the model learns to predict the next word during pre-training with massive text data. Then, during the fine-tuning phase, it is optimized for dialogue or question-answering using human-labeled data. Many models are guided to provide safer, more user-friendly responses through \u0026ldquo;Reinforcement Learning from Human Feedback\u0026rdquo;. When a user asks a question, the model predicts the most appropriate response based on context and statistical probabilities (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eSome users may find AI chatbots appealing due to the informative and accessible content they provide. However, debates continue regarding the reliability and accuracy of AI chatbots. AI chatbots may sometimes exhibit a tendency to \u0026ldquo;hallucinate\u0026rdquo; (generate unrealistic or incorrect information) (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Some studies highlight the potential of chatbots to provide inaccurate information when generating medical content (\u003cspan additionalcitationids=\"CR15 CR16\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). Previous studies in several specialties of dentistry have also determined that AI chatbots can provide users with incorrect information (\u003cspan additionalcitationids=\"CR18 CR19 CR20\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). In addition, it has been reported that the \u0026ldquo;Reinforcement Learning from Human Feedback\u0026rdquo; process does not completely eliminate the risk of the model being manipulated or misdirected (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). The potential to generate inaccurate, outdated, or biased content, particularly in health information, is a critical concern for patient safety (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eWhether patients obtain information from forums or AI chatbots, the readability of the text is as important as its accuracy in enabling patients to draw the necessary conclusions. The readability of a text is found by its Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Readability tools score the text by assessing the complexity of vocabulary and syntax. However, they do not directly measure understandability and actionability (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). The Patient Education Materials Assessment Tool (PEMAT-P) was designed to assess the understandability of a patient information/education text. PEMAT-P also has an additional seven questions, which are used to evaluate actionability (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eResearch on the accuracy, quality, comprehensibility, and applicability of AI chatbots in answering dentistry-related questions has become an increasingly common and actively explored topic. (\u003cspan additionalcitationids=\"CR19 CR20\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan additionalcitationids=\"CR24 CR25 CR26 CR27\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). However, the varying results reported across these studies make it difficult to generalize that any AI chatbot can consistently provide high-quality responses across all topics. Therefore, it remains uncertain whether AI chatbots can reliably serve as comprehensive tools for patient education. Additionally, to the best of our knowledge, there are no studies evaluating the performance of AI chatbots in smile design. Thus, this study aimed to assess and compare the accuracy and quality of the information provided by five AI chatbots to patient questions. Additionally, the readability, understandability, and actionability of responses of chatbots were also evaluated. The null hypotheses of this study were that (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) accuracy and quality of responses would not differ among five different AI chatbots nor existing literature, and (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) readability, understandability, and actionability of responses would not differ.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eIn this study, the methodology of a previous study was used to determine patient concerns and questions regarding smile design and the method (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Searches were conducted on the popular question-and-answer sites Reddit (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.reddit.com/r/asksdentists/\u003c/span\u003e\u003cspan address=\"https://www.reddit.com/r/asksdentists/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and Quora (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.quora.com/\u003c/span\u003e\u003cspan address=\"https://www.quora.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) using the search keywords \u0026ldquo;smile design\u0026rdquo; and \u0026ldquo;anterior esthetic dental restorations\u0026rdquo;. Questions were recorded. Using these records, a total of 28 questions were prepared under the headings of general questions, treatment procedure, care and maintenance, and functional and aesthetic concerns (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). The questions were prepared using as little dental terminology as possible, so that they appear to be asked by the patient rather than a dental professional (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eDefined questions about smile design and anterior dental esthetics\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"1\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eGeneral Questions\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e1. Which restoration type is best for my front tooth: direct laminate, indirect laminate, or full crown?\u003c/p\u003e\u003cp\u003e2. What happens if I have short teeth or a \u0026ldquo;gummy smile\u0026rdquo;?\u003c/p\u003e\u003cp\u003e3. Which materials give the most natural result \u0026mdash; metal-ceramic, full ceramic, or composite?\u003c/p\u003e\u003cp\u003e4. Are these cosmetic gum and tooth treatments permanent or reversible?\u003c/p\u003e\u003cp\u003e5. How much will this kind of smile design cost, and how long will it last?\u003c/p\u003e\u003cp\u003e6. Is gum health important before doing veneers or crowns?\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eTreatment \u0026amp; Procedure\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e7. What exactly happens during a gum contouring?\u003c/p\u003e\u003cp\u003e8. How long does it take for gums to heal after reshaping?\u003c/p\u003e\u003cp\u003e9. What if the gum surgery is done incorrectly \u0026mdash; can it be fixed later?\u003c/p\u003e\u003cp\u003e10. Will gum surgery change the way my teeth look immediately?\u003c/p\u003e\u003cp\u003e11. For laminates and full crowns, how much tooth needs to be removed?\u003c/p\u003e\u003cp\u003e12. How many appointments are needed for veneers and gum reshaping?\u003c/p\u003e\u003cp\u003e13. How do dentists decide where my new gum line should be?\u003c/p\u003e\u003cp\u003e14. If my gums grow back or shrink after the surgery, will it affect my restorations?\u003c/p\u003e\u003cp\u003e15. Can I combine gum reshaping and laminate placement in the same visit?\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eCare \u0026amp; Maintenance\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e16. How can I prevent my gums from becoming inflamed again after cosmetic treatment?\u003c/p\u003e\u003cp\u003e17. Do veneers or crowns require special flossing techniques around the gumline?\u003c/p\u003e\u003cp\u003e18. How often should I have professional cleaning or maintenance after smile design?\u003c/p\u003e\u003cp\u003e19. Can gums recede after veneers or crowns, and what can I do to prevent that?\u003c/p\u003e\u003cp\u003e20. Can whitening or mouthwash irritate gums after cosmetic procedures?\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eFunctional And Aesthetic Concerns\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e21. How much do gums affect the beauty of my smile?\u003c/p\u003e\u003cp\u003e22. Will reshaping my gums make my teeth look longer and more even?\u003c/p\u003e\u003cp\u003e23. What if the gum levels become uneven \u0026mdash; can it ruin the symmetry of my smile?\u003c/p\u003e\u003cp\u003e24. Can poorly shaped gums make veneers or crowns look fake?\u003c/p\u003e\u003cp\u003e25. Will gum surgery change how my lips move when I smile?\u003c/p\u003e\u003cp\u003e26. Can gum color be adjusted or treated if it looks uneven after surgery?\u003c/p\u003e\u003cp\u003e27. How long will it take for my gums to look completely natural again?\u003c/p\u003e\u003cp\u003e28. Does gum surgery have any risk of affecting tooth sensitivity or root exposure?\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eIn this study, since no human subjects or personal data were used, ethical approval was not needed in accordance with the Declaration of Helsinki. The research involved only the evaluation of AI chatbot responses to pre-defined questions, ensuring that no identifiable personal information was collected or analyzed.\u003c/p\u003e\u003cp\u003eA volunteer who did not participate in the study asked the questions to the Microsoft Copilot, Google Gemini, ChatGPT-4, ChatGPT-5, and Claude Sonet 4.5 chatbots in the same order and on the same day on November 10, 2025. To prevent the answer from affecting earlier questions, a new chat window was opened for each question. All responses were saved in a Microsoft Word file with the original formatting preserved. Each chatbot\u0026rsquo;s responses document was number-coded before being handled to authors for blinded assessment. For blinded assessment, the volunteer also prepared a document for the evaluation of each question (See Supplementary file 1).\u003c/p\u003e\u003cp\u003eTwo prosthodontists (authors) rated all responses\u0026rsquo; accuracy, quality, understandability, and actionability separately. For accuracy assessment, a five-point Likert scale ranging from 1 (the chatbot\u0026rsquo;s answer is completely incorrect), to 2 (the chatbot\u0026rsquo;s answer contains more incorrect items than correct items), to 3 (the chatbot\u0026rsquo;s answer contains an equal balance of correct and incorrect items), to 4 (the chatbot\u0026rsquo;s answer contains more correct items than incorrect items), and to 5 (the chatbot\u0026rsquo;s answer is completely correct) was used (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). The Global Quality Scale (GQS) with a 5-point scoring system was used to evaluate the quality of the responses (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eFRE and FKGL scores were automatically calculated using a freely available online tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://goodcalculators.com/flesch-kincaid-calculator/\u003c/span\u003e\u003cspan address=\"https://goodcalculators.com/flesch-kincaid-calculator/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). This online tool calculates FRE and FKGL using the following formulas:\u003c/p\u003e\u003cp\u003eFRE score\u0026thinsp;=\u0026thinsp;206.835\u0026thinsp;\u0026minus;\u0026thinsp;1.015 (total number of words/total number of sentences)\u0026thinsp;\u0026minus;\u0026thinsp;84.6 (total number of syllables/total number of words).\u003c/p\u003e\u003cp\u003eFKGL score\u0026thinsp;=\u0026thinsp;0.39 (total number of words/total number of sentences)\u0026thinsp;+\u0026thinsp;11.8 (total number of syllables/total number of words)\u0026thinsp;\u0026minus;\u0026thinsp;15.59.\u003c/p\u003e\u003cp\u003eFRE scores can be interpreted by referencing established conversion tables to estimate the reader\u0026rsquo;s required education level, whereas FKGL directly reports the minimum grade level needed to comprehend the text (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e).\u003c/p\u003e\u003cp\u003ePEMAT-P was used to assess the understandability and actionability of responses of AI chatbots. The PEMAT-P consists of 17 questions for understandability and 7 questions for actionability. First, prosthodontists read the PEMAT-P manual and examined examples related to scoring. Each author independently rated all responses as 0 for \u0026ldquo;disagree\u0026rdquo;, 1 for \u0026ldquo;agree\u0026rdquo;, and N/A for \u0026ldquo;not applicable\u0026rdquo; (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Since the responses did not have visual aids, the relevant questions of PEMAT-P were marked \u0026ldquo;not applicable\u0026rdquo; and excluded from the evaluation (See Supplementary file 2).\u003c/p\u003e\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eStatistical analysis\u003c/h2\u003e\u003cp\u003eData were statistically analyzed using a statistical software (SPSS, version 22.0, IBM Corp., Armonk, NY). Cohen\u0026rsquo;s kappa was calculated to evaluate the inter-rater agreement between the two researchers. Inter-rater agreement for the 5-point Likert scale, GQS, PEMAT-P/Understandability, and PEMAT-P/Actionability was assessed using the Cohen\u0026rsquo;s kappa statistic. Cohen\u0026rsquo;s kappa values of 0.00-0.20 were considered slight agreement, 0.21\u0026ndash;0.40 fair, 0.41\u0026ndash;0.60 moderate, 0.61\u0026ndash;0.80 substantial, and 0.81-1.00 almost perfect. Data normality was checked using the Kolmogorov-Smirnov test, and Levene\u0026rsquo;s test was used to assess the homogeneity of the variances. 5-point Likert scale, GQS, PEMAT-P/Understandability, and PEMAT-P/Actionability were non-normally distributed. These variables were analyzed using the Kruskal-Wallis test. Differences among the AI chatbot responses were determined using Dunn\u0026rsquo;s test. FRE and FKGL scores, normally distributed, were analyzed using a one-way analysis of variance (ANOVA) and Tukey\u0026rsquo;s HSD post hoc tests. Statistical significance was set at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e\u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eCohen\u0026rsquo;s kappa values indicated consistent inter-rater reliability across all measures. The CQS showed moderate agreement (κ\u0026thinsp;=\u0026thinsp;0.55), while the PEMAT-P understandability demonstrated substantial agreement (κ\u0026thinsp;=\u0026thinsp;0.73). Both the 5-point Likert scale and PEMAT-P actionability assessments showed almost perfect agreement (κ\u0026thinsp;=\u0026thinsp;0.85 and κ\u0026thinsp;=\u0026thinsp;0.84, respectively). The results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison of Accuracy, Quality, FRE, FKGL, PEMAT-P/Understandability, and PEMAT-P/Actionability scores according to AI chatbots.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMicrosoft Copilot\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eChatGPT-4\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eGoogle Gemini\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eChatGPT-5\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eClaude Sonet 4.5\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003ep\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003cp\u003eFive-point Likert scale\u003c/p\u003e\u003cp\u003eMedian (min-max)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5 (\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5 (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5 (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e5 (\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e5 (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e0. 013\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eQuality\u003c/p\u003e\u003cp\u003eGQS scores\u003c/p\u003e\u003cp\u003eMedian\u003c/p\u003e\u003cp\u003e(min-max)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e5 (\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003eac\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e5 (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5 (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e4 (\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3 (\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e\u0026lt;\u0026thinsp;0.001\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFRE\u003c/p\u003e\u003cp\u003eMean\u003c/p\u003e\u003cp\u003e(Standard deviation)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e43.26 (10.6)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e41.27 (10.83)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e49.05 (6.59)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e37.33 (11.95)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e42.26 (14.25)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e0.004\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFKGL\u003c/p\u003e\u003cp\u003eMean\u003c/p\u003e\u003cp\u003e(Standard deviation)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e9.02 (1.55)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e9.27 (1.44)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e9.76 (1.03)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e12.78 (1.98)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e8.52 (1.97)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e\u0026lt;\u0026thinsp;0.001\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePEMAT-P Understandability\u003c/p\u003e\u003cp\u003eMedian\u003c/p\u003e\u003cp\u003e(min-max)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e84.61 (71.42\u0026ndash;100)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e84.61 (57.14\u0026ndash;92.85)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e83.33 (58.33\u0026ndash;92.85)\u003csup\u003eab\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e75 (45.45- 90)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e83.33 (42.85\u0026ndash;92.3)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e\u0026lt;\u0026thinsp;0.001\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePEMAT-P Actionability\u003c/p\u003e\u003cp\u003eMedian\u003c/p\u003e\u003cp\u003e(min-max)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e40 (0\u0026ndash;80)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e20 (0\u0026ndash;80)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e40(0\u0026ndash;60)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e20 (0\u0026ndash;60)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e40(0\u0026ndash;80)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e\u0026lt;\u0026thinsp;0.001\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"7\"\u003e(GQS: Global Quality Scale, FRE: Flesch Reading Ease, FKGL: Flesch-Kincaid Grade Level, PEMAT-P: Patient Education Materials Assessment Tool)\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eA statistically significant difference was found in the median accuracy scores among AI chatbots, \u003cem\u003eX\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e (4, 280)\u0026thinsp;=\u0026thinsp;12.658, p\u0026thinsp;=\u0026thinsp;0. 013. The accuracy scores of ChatGPT-5 were lower than both Google Gemini and Claude Sonet 4.5 (p\u0026thinsp;=\u0026thinsp;0.017 and p\u0026thinsp;=\u0026thinsp;0.041, respectively). A statistically significant difference was also found in the median quality scores among AI chatbots, \u003cem\u003eX\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e (4, 280)\u0026thinsp;=\u0026thinsp;104.525, p\u0026thinsp;=\u0026thinsp;0.001. The quality scores of ChatGPT-5 were significantly different from all other AI chatbots (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). There were also significant differences between the quality scores of Claude Sonet 4.5 and both ChatGPT-4 (p\u0026thinsp;=\u0026thinsp;0.042) and Google Gemini (p\u0026thinsp;=\u0026thinsp;0.025).\u003c/p\u003e\u003cp\u003eOne-way ANOVA revealed significant differences in FRE scores of AI chatbot responses (F(4,135)\u0026thinsp;=\u0026thinsp;4.047, p\u0026thinsp;=\u0026thinsp;0.004). Tukey\u0026rsquo;s HSD Test for multiple comparisons found that the mean FRE scores were significantly different between Google Gemini and ChatGPT-5 (p\u0026thinsp;=\u0026thinsp;0.001). There was no statistically significant difference in mean FRE scores between other AI chatbots (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). One-way ANOVA also revealed significant differences in FKGL scores of AI chatbot responses (F(4,135)\u0026thinsp;=\u0026thinsp;30.838, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Higher FKGL scores were calculated in ChatGPT-5 than other AI chatbots (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). FKGL scores of Google Gemini were higher than those of Claude Sonet 4.5 (p\u0026thinsp;=\u0026thinsp;0.043). No significant difference was found between other AI chatbots (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05). As listed in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the majority of AI chatbot responses were written at an 8th-grade level or higher. All ChatGPT-5 responses corresponded to an education level of 8th grade or higher. The proportions of responses below an 8th-grade level were 25% for Microsoft Copilot, 17.85% for ChatGPT-4, 3.6% for Google Gemini, and 42.85% for Claude Sonet 4.5.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eNumber and percentage of the Flesch-Kincaid Grade Level (FKGL) of the AI chatbot responses\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFKGL (Readability Level)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMicrosoft Copilot\u003c/p\u003e\u003cp\u003en (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eChatGPT-4\u003c/p\u003e\u003cp\u003en (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eGoogle Gemini\u003c/p\u003e\u003cp\u003en (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eChatGPT-5\u003c/p\u003e\u003cp\u003en (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eClaude Sonet 4.5\u003c/p\u003e\u003cp\u003en (%)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e5 (Very Easy)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e6 (Easy)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e3 (10.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e8 (28.6)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e7 (Fairly Easy)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e4 (14.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4 (14.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e4 (14.3)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e8\u0026ndash;9 (Standard)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e13 (46.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e14 (50)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e16 (57.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e3 (10.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e10 (35.7)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e10\u0026ndash;12 (Fairly Difficult)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e7 (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e9 (32.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e11 (39.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e12 (42.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e6 (21.4)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e13\u0026ndash;16 (Difficult)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1 (3.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e11 (39.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCollege Graduate (Very Difficult)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2 (7.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0 (0)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eFigure \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the distribution of PEMAT-P understandability and actionability scores. Kruskal\u0026ndash;Wallis analysis showed that PEMAT-P understandability scores of AI chatbots were significantly different, \u003cem\u003eX\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e (4, 280)\u0026thinsp;=\u0026thinsp;21.726, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001. The median PEMAT-P/Understandability score of Microsoft Copilot was significantly different than ChatGPT-5 (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and Claude Sonet 4.5 (p\u0026thinsp;=\u0026thinsp;0.016). Kruskal\u0026ndash;Wallis analysis also revealed that the PEMAT-P Actionability scores of AI chatbots were significantly different, \u003cem\u003eX\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e (4, 280)\u0026thinsp;=\u0026thinsp;46.727, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001. Dunn\u0026rsquo;s pairwise comparisons revealed that the PEMAT-P Actionability scores of ChatGPT-5 were significantly different from the PEMAT-P Actionability scores of other AI chatbots.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study compared the performance of five different AI chatbots using multiple evaluation tools, including the 5-point Likert scale, the GQS, PEMAT-P Understandability, PEMAT-P Actionability, and readability indices (FRE and FKGL). The results suggested that chatbot-generated health information shows considerable variability. Overall, Microsoft Copilot, ChatGPT-4, and Google Gemini tended to provide higher-quality and more comprehensible responses, while ChatGPT-5 and Claude Sonet 4.5 showed lower performance in accuracy and quality. Therefore, the first null hypothesis was rejected. The readability, understandability, and actionability of AI chatbot responses were also significantly different. The second null hypothesis was also rejected.\u003c/p\u003e\u003cp\u003eIn multi-step smile design procedures that include surgical interventions, the accuracy and quality of patient information help ensure that patients correctly understand the treatment process, which in turn supports improved treatment outcomes. Whether the information being accurate or of poor quality can directly influence the patient\u0026rsquo;s understanding, expectations, and overall satisfaction with the treatment. Therefore, the responses provided by AI chatbots, which patients may easily use to obtain information about the procedure, should also be accurate and of high quality. However, in this study, the significant differences observed in the 5-point Likert scale and GQS ratings reflect variations in the overall reliability, completeness, and usefulness of the AI chatbot responses. Higher-scoring models appeared to provide more structured explanations, fewer irrelevant details, and better alignment with established clinical knowledge. The high accuracy and quality of the Microsoft Copilot, ChatGPT-4, and Google Gemini observed in this study are consistent with those reported in a previous study (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). In contrast, lower-scoring AI chatbots (ChatGPT-5 and Claude Sonet 4.5) often demonstrated oversimplification, occasional inconsistencies, and limited contextual depth. When patients are not accurately informed, they may develop unfounded fears or concerns about the safety of smile design procedures, which may cause them to hesitate or refuse to undergo the necessary procedures. Additionally, conflicting information between AI chatbots and dentists can undermine trust and prevent patients from following professional advice or seeking timely treatment for any future dental procedures (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eSome studies have reported that the scientific accuracy of AI chatbot responses increases depending on the sophistication of the Natural Language Processing (NLP) system and the use of larger datasets (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e). Despite their advanced NLP architectures and large training datasets, several reasons may explain why ChatGPT-5 and Claude Sonet 4.5 show reduced accuracy and quality. The responses generated by ChatGPT-5 were excessively brief and lacked the level of informational depth expected even for patient-oriented content. This response pattern may have resulted from several factors, including safety filters and alignment mechanisms. The safety filters and alignment mechanisms of ChatGPT-5 can make their answers more conservative and less specific, which lowers perceived quality in complex or technical domains (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). On the other hand, although Claude Sonet 4.5 gave long answers, some of them were scientifically incorrect. This is likely because the model can produce fluent but wrong information (hallucinations), which may seem convincing even when it is not accurate (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). Consistent with the results of this study, Tuzlalı et al.(\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e) also mentioned the limited accuracy and completeness of Claude responses for questions about dental implants. This shows that there is a risk of errors or missing information in some responses. Improving the suitability of an AI model for dentistry topics can be achieved by providing additional training to the base model. This process creates a dentistry-focused expert model and may significantly increase its accuracy (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAn accurate text with high readability can help patients understand medical information more quickly, participate more actively in decision-making, and follow treatment procedures more accurately. Patient education materials should be written at or below the 8th grade-level (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). Materials written above this level may reduce comprehension and limit effective patient engagement. In this study, within these criteria, the lowest readability scores were observed for ChatGPT-5 (with an FRE of 37.33 and an FKGL of 12.78). The high FKGL scores of ChatGPT‑5 responses may be attributed to the use of technical terminology and complex sentence structures, which increase reading difficulty (\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eThe five AI chatbots evaluated in this study produced responses with mean FRE scores between 37.33 and 49.05 and mean FKGL scores between 8.52 and 12.78. When evaluating the FRE scores, no statistically significant differences were observed among the AI chatbots except for ChatGPT‑5, whereas the lowest FKGL scores were found in Microsoft Copilot, ChatGPT‑4, and Claude Sonet 4.5. In the literature on dental procedures, the readability scores of AI chatbot responses vary across studies. \u0026Ouml;zcivelek and \u0026Ouml;zcan (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e), in their study on maxillofacial prosthesis, reported that the FKGL scores of the four AI chatbots (DeepSeek-R1, Chat GPT-o1, ChatGPT-4, and Dental GPT) varied between 9.42 and 10.70, corresponding to a \u0026lsquo;fairly difficult\u0026rsquo; reading level. \u0026Ouml;zdemir et al. (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e), in their study on oral habits, reported average FKGL scores of 10.47 for ChatGPT‑4, 9.53 for Google Gemini, and 9.61 for Microsoft Copilot. ChatGPT‑4 responses were found to be more complex and less readable than those of Gemini and Copilot. Esmailpour et al.(\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e) also reported that ChatGPT-4 (10.45) produced higher FKGL scores compared to Microsoft Copilot (8.38) and Google Gemini (7.82).In contrast to \u0026Ouml;zdemir et al. and Esmailpour et al, our study found no significant differences in FKGL scores between ChatGPT‑4, Microsoft Copilot, and Google Gemini. G\u0026uuml;ven et al.(\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e), in their study on traumatic dental injuries, also found no significant difference between ChatGPT-4 and Google Gemini, although both produced responses at a college-level readability. In a study evaluating responses on endodontic iatrogenic events of four AI chatbots (ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4), the authors reported that the generated answers were classified as \u0026lsquo;fairly difficult\u0026rsquo; (high school) to \u0026lsquo;difficult\u0026rsquo; (college) to read (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e). As seen in previous studies, when the same AI chatbots are asked questions on different dental topics, the readability of their responses can vary considerably. It is due to questions about specific procedures or technical topics making the AI chatbot use more complex words or longer sentences, which raise the reading level, while general patient-focused questions produce simpler and easier-to-read text (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e). In addition, adding prompts that include simple language rules (short sentences, common words, active sentence structure) to the questions can improve readability (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e). However, individuals who do not routinely use AI chatbots and who are not aware of their limitations may not always be able to apply such prompts to improve the readability of the information provided.\u003c/p\u003e\u003cp\u003eReadability alone may not ensure that patients understand dental procedures or can make informed decisions. Therefore, assessing the responses using PEMAT‑P is needed to evaluate the understandability and actionability. In the present study, ChatGPT-5 showed the lowest PEMAT-P understandability scores (75%), which was consistent with its poorer readability results. The median of PEMAT-P understandability scores of the other AI chatbots ranged from 83.33% to 84.61%. The literature recommends a 70% threshold for both PEMAT-P understandability and actionability (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e). The PEMAT-P user guide also states that materials scoring above this level can be interpreted as \u0026ldquo;more understandable and more actionable. In previous studies (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan additionalcitationids=\"CR38\" citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e), most AI chatbots show PEMAT-P understandability scores of 70% or higher for patient-education responses related to dental topics. In this study, consistent with the literature, most of the understandability scores of AI chatbot responses were found to be above the 70% threshold (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). However, most actionability scores were found to be below this threshold. Additionally, ChatGPT-5 was found to have lower actionability scores compared to the other AI chatbots (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Several studies (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan additionalcitationids=\"CR38\" citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e) in the literature support our finding that AI chatbot responses show high understandability but low actionability. The reason for lower actionability of AI-generated patient materials may be due to the use of medical terminology and general, non-guiding statements (for example, step-by-step action plans, checklists, and clear recommendations) aimed at preventing errors (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eIn this study, the accuracy and quality of information provided to patients by AI chatbots in complex dental procedures involving multiple procedures, such as smile design, were evaluated. As in other studies, significant differences were found between the AI chatbots in this study. In this study, the identification of the lowest accuracy and quality values in ChatGPT-5 and Claude Sonnet 4.5 reveals that some models still have significant limitations in complex dental decision-making processes. This study has also revealed that factors such as the readability and actionability of AI chatbot responses may limit their use in patient education. However, there are several limitations that may affect the results of this study. First, the study was designed so that a new chat window was used for each question. Each question is entered separately without allowing follow-up questions or explanations, which naturally limits the depth and realism of chatbot interactions. This approach did not assess chatbots' ability to improve, contextualize, and elaborate responses through multi-turn dialogues. Additionally, preparing questions in a structured template format (diagnosis \u0026rarr; explanation \u0026rarr; options \u0026rarr; follow-up) can also improve the response performance of AI chatbots. However, this method was not preferred by the authors, considering the limited intellectual potential of an average individual to use the chatbot in this sequence. Additionally, for questions related to clinical scenarios, criteria such as accuracy and quality may be influenced by the individual judgment of the authors, which can affect the assessment results. In our study, 5-point Likert and Global Quality Scales were used, and high inter-rater reliability results were found on the Likert scale (κ\u0026thinsp;=\u0026thinsp;0.85) and moderate results on the GQS (κ\u0026thinsp;=\u0026thinsp;0.55). This situation shows that a certain degree of variability is inevitable in human evaluations. Hybrid evaluation models that combine expert assessments with AI-powered content scoring models can increase consistency and reduce bias (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). Furthermore, the study evaluated only five publicly available chatbots as of November 2025. The performance and capabilities of AI chatbots are rapidly improving with model updates and fine-tuning iterations. This situation is one of the reasons why studies conducted at different time intervals yield different results. Another limitation of the study is that it was conducted only in English. AI chatbot responses in different languages may differ in terms of performance.\u003c/p\u003e\u003cp\u003eAI applications have great potential to develop more accessible and readable online content for patient education. In addition, they can be a step towards increasing patient satisfaction through education. Future studies may evaluate artificial intelligence technologies specialized in dental topics that can create visual aids to answer patient questions.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study assessed the accuracy, quality, readability, understandability, and actionability of responses from five AI chatbots for dental questions, including complex procedures like smile design. ChatGPT-5 and Claude Sonet 4.5 had the lowest accuracy and quality scores, which show limits in handling multi-step dental cases. Most responses were understandable, but their readability often needed a high reading level, which may make them harder for general patients to follow. Actionability scores were generally low.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAI: Artificial intelligence\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLLM: Large Language Model\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFRE: Flesch Reading Ease\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFKGL: Flesch-Kincaid Grade Level\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePEMAT-P: Patient Education Materials Assessment Tool\u003c/p\u003e\n\u003cp\u003eGQS: Global Quality Scale\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eANOVA: One-way analysis of variance\u003c/p\u003e\n\u003cp\u003eNLP: Natural Language Processing\u0026nbsp;\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study did not require ethical approval in accordance with the Declaration of Helsinki, as it involved only evaluation of AI chatbot responses and did not include human subjects or personal data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: B.E. and H.A.U.; Methodology: B.E. and H.A.U.; Formal analysis and investigation B.E. and H.A.U.; Writing\u0026mdash;original draft preparation: B.E. and H.A.U.; Writing\u0026mdash;review and editing: B.E.; Funding acquisition: B.E and H.A.U.; Laboratory resources: - Supervision: B.E.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to thank Dr. Berna Kavuncu Ezmek for her contributions in obtaining and editing the AI chatbot responses.\u003c/p\u003e\n"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAfroz S, Rathi S, Rajput G, Rahman SA. Dental esthetics and its impact on psycho-social well-being and dental self confidence: A campus based survey of north indian university students. J Indian Prosthodont Soc. 2013;13(4):455\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAbbasi MS, Lal A, Das G, Salman F, Akram A, Ahmed AR, et al. Impact of Social Media on Aesthetic Dentistry: General Practitioners\u0026rsquo; Perspectives. Healthc. 2022;10(10):1\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBaik KM, Anbar G, Alshaikh A, Banjar A. Effect of Social Media on Patient\u0026rsquo;s Perception of Dental Aesthetics in Saudi Arabia. Int J Dent. 2022;2022:4794497.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStojilković M, Gušić I, Berić J, Prodanović D, Pecikozić N, Veljović T, et al. Evaluating the influence of dental aesthetics on psychosocial well-being and self-esteem among students of the University of Novi Sad, Serbia: a cross-sectional study. BMC Oral Health. 2024;24(1):1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eArmalaite J, Jarutiene M, Vasiliauskas A, Sidlauskas A, Svalkauskiene V, Sidlauskas M, et al. Smile aesthetics as perceived by dental students: A cross-sectional study. BMC Oral Health. 2018;18(1):1\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMariam A, Prakash VS, Graduate P. Digital Smile Design: Revolutionizing Aesthetic Dentistry Through Technological Innovation. IJSDR. 2024;9(6):1187\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDoğan AN. Dijital G\u0026uuml;l\u0026uuml;ş Tasarimi: KullanilanSi̇stemler Ve Avantajlari. Sağlık Bilim Derg. 2020;29(2):138\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAktan E. Diş Hekimliğinde Dijital G\u0026uuml;l\u0026uuml;ş Tasarımı Uygulamaları Digital Smile Design Applications in Dentistry. ADO J Clin Sci. 2023;12(3):474\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYun HS, Bickmore T. Online Health Information-Seeking in the Era of Large Language Models: Cross-Sectional Web-Based Survey Study. J Med Internet Res. 2025;27.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZeng L, Li Q, Zuo Y, Zhang Y, Li Z. Perceptions and Attitudes of Chinese Oncologists Toward Endorsing AI-Driven Chatbots for Health Information Seeking Among Patients with Cancer: Phenomenological Qualitative Study. J Med Internet Res. 2025;27:1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBridgelall R. Unraveling the mysteries of AI chatbots. Artif Intell Rev. 2024;57:89.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGonz\u0026aacute;lez Barman K, Lohse S, de Regt HW. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values. Whose Perspectives? Philos Technol. 2025;38:35.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChow JCL, Li K. Large Language Models in Medical Chatbots: Opportunities, Challenges, and the Need to Address AI Risks. Inf. 2025;16(7):1\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLawson McLean A, Hristidis V. Evidence-Based Analysis of AI Chatbots in Oncology Patient Education: Implications for Trust, Perceived Realness, and Misinformation Management. J Cancer Educ. 2025;40(4):482\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePeykani P, Ramezanlou F, Tanasescu C. Large Language Models: A Structured Taxonomy and Review of Challenges, Limitations, Solutions, and Future Directions. Appl Sci. 2025;15:8103.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang W, Zhang J. Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics. 2025;13:856.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTuzlalı M, Baki N, Aral K, Aral CA. Evaluating the performance of AI chatbots in responding to dental implant FAQs: A comparative study. BMC Oral Health. 2025;25:1548.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e\u0026Ouml;zcivelek T, \u0026Ouml;zcan B. Comparative evaluation of responses from DeepSeek-R1, ChatGPT-o1, ChatGPT-4, and dental GPT chatbots to patient inquiries about dental and maxillofacial prostheses. BMC Oral Health. 2025;25:871.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHelvacioglu-Yigit D, Demirturk H, Ali K, Tamimi D, Koenig L, Almashraqi A. Evaluating artificial intelligence chatbots for patient education in oral and maxillofacial radiology. Oral Surg Oral Med Oral Pathol Oral Radiol. 2025;139(6):750\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eEsmailpour H, Rasaie V, Babaee Hemmati Y, Falahchai M. Performance of artificial intelligence chatbots in responding to the frequently asked questions of patients regarding dental prostheses. BMC Oral Health. 2025;25:574.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGuven Y, Ozdemir OT, Kavan MY. Performance of Artificial Intelligence Chatbots in Responding to Patient Queries Related to Traumatic Dental Injuries: A Comparative Study. Dent Traumatol. 2025;41(3):338\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDahlgren Lindstr\u0026ouml;m A, Methnani L, Krause L, Ericson P, de Rituerto de Troya \u0026Iacute;M, Coelho Mollo D, et al. Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. Ethics Inf Technol. 2025;27(2):1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMohammad-Rahimi H, Ourang SA, Pourhoseingholi MA, Dianat O, Dummer PMH, Nosrat A. Validity and reliability of artificial intelligence chatbots as public sources of information on endodontics. Int Endod J. 2024;57(3):305\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFreire Y, Santamar\u0026iacute;a Laorden A, Orejas P\u0026eacute;rez J, G\u0026oacute;mez S\u0026aacute;nchez M, D\u0026iacute;az-Flores Garc\u0026iacute;a V, Su\u0026aacute;rez A. ChatGPT performance in prosthodontics: Assessment of accuracy and repeatability in answer generation. J Prosthet Dent. 2024;131(4):659.e1-659.e6.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBabayiğit O, Tastan Eroglu Z, Ozkan Sen D, Ucan Yarkac F. Potential Use of ChatGPT for Patient Information in Periodontology: A Descriptive Pilot Study. Cureus. 2023;15(11):e48518.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDaraqel B, Wafaie K, Mohammed H, Cao L, Mheissen S, Liu Y et al. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop [Internet]. 2024;165(6):652\u0026ndash;62. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://linkinghub.elsevier.com/retrieve/pii/S0889540624000593\u003c/span\u003e\u003cspan address=\"https://linkinghub.elsevier.com/retrieve/pii/S0889540624000593\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAguiar de Sousa R, Costa SM, Almeida Figueiredo PH, Camargos CR, Ribeiro BC, Alves e Silva MRM. Is ChatGPT a reliable source of scientific information regarding third-molar surgery? J Am Dent Assoc [Internet]. 2024;155(3):227\u0026ndash;232.e6. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://linkinghub.elsevier.com/retrieve/pii/S0002817723006815\u003c/span\u003e\u003cspan address=\"https://linkinghub.elsevier.com/retrieve/pii/S0002817723006815\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKılın\u0026ccedil; DD, Mansız D. Examination of the reliability and readability of Chatbot Generative Pretrained Transformer\u0026rsquo;s (ChatGPT) responses to questions about orthodontics and the evolution of these responses in an updated version. Am J Orthod Dentofac Orthop [Internet]. 2024;165(5):546\u0026ndash;55. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://linkinghub.elsevier.com/retrieve/pii/S0889540624000076\u003c/span\u003e\u003cspan address=\"https://linkinghub.elsevier.com/retrieve/pii/S0889540624000076\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYagci F, Eraslan R, Albayrak H, İpekten F. Accuracy and Reliability of Artificial Intelligence Chatbots as Public Information Sources in Implant Dentistry. Int J Oral Maxillofac Implants. 2025;1\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAzadi A, Gorjinejad,Fatemeh Mohammad-Rahimi H, Tabrizi R, Alam M, Golkar M. Evaluation of AI-generated responses by different artificial intelligence chatbots to the clinical decision-making case-based questions in oral and maxillofacial surgery. Oral Surg Oral Med Oral Pathol Oral Radiol. 2024;137(6):587\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAsfuroğlu ZM, Yağar H, G\u0026uuml;m\u0026uuml;şoğlu E. High accuracy but limited readability of large language model-generated responses to frequently asked questions about Kienb\u0026ouml;ck\u0026rsquo;s disease. BMC Musculoskelet Disord. 2024;25:879.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e\u0026Ouml;zdemir \u0026Ouml;T, Kavan MY, G\u0026uuml;ven Y. Evaluation of the readability, quality, and accuracy of AI chatbot responses to questions about deleterious oral habits. BMC Oral Health. 2025;25:1812.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTaşy\u0026uuml;rek M, Adıg\u0026uuml;zel \u0026Ouml;, Orta\u0026ccedil; H. Comparative Evaluation of Responses from ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4 Chatbots to Questions About Endodontic Iatrogenic Events. Healthc. 2025;13:2615.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOkuhara T, Furukawa E, Okada H, Yokota R, Kiuchi T. Readability of written information for patients across 30 years: A systematic review of systematic reviews. Patient Educ Couns [Internet]. 2025;135(October 2024):108656. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.pec.2025.108656\u003c/span\u003e\u003cspan address=\"10.1016/j.pec.2025.108656\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSivaramakrishnan G, Almuqahwi M, Ansari S, Lubbad M, Alagamawy E, Sridharan K. Assessing the power of AI: a comparative evaluation of large language models in generating patient education materials in dentistry. BDJ Open. 2025;11(1):1\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShoemaker SJ, Wolf MSBC. Developement of the Patient Education Material Assessment Tool (PEMAT). Physiol Behav. 2018;176(5):139\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSpuur K, Currie G, Al-Mousa D, Pape R. Suitability of ChatGPT as a Source of Patient Information for Screening Mammography. Health Promot Pract. 2025;26(4):746\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAlnsour MM, Alenezi R, Barakat M, AL-Omiri MK. Assessing ChatGPT\u0026rsquo;s suitability in responding to the public\u0026rsquo;s inquires on the effects of smoking on oral health. BMC Oral Health. 2025;25(1).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHalboub E, Hakami RAM, Khalufi KNA, Hakami SAH, Alhajj MN. Assessment of quality, understandability, actionability, and readability of responses of selected chatbots to the top searched queries about oral cancer. Digit Dent J. 2025;1:100008.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eS\u0026ouml;nmezoğlu Hİ, G\u0026uuml;ner S\u0026ouml;nmezoğlu B, Temel MH, \u0026Ccedil;akir B. Comprehensibility and readability of selected artificial intelligence chatbots in providing uveitis-related information. Med (Baltim). 2025;104(43):e45135.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"AI chatbots, smile design, accuracy assessment, health information quality, readability, understandability and actionability","lastPublishedDoi":"10.21203/rs.3.rs-8210813/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8210813/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground:\u003c/h2\u003e\u003cp\u003eThis study aimed to evaluate and compare the accuracy, quality, readability, understandability, and actionability of responses provided by five AI chatbots\u0026mdash;Microsoft Copilot, ChatGPT-4, ChatGPT-5, Google Gemini, and Claude Sonet 4.5\u0026mdash;to patient questions about smile design and anterior aesthetic dental procedures.\u003c/p\u003e\u003ch2\u003eMethod:\u003c/h2\u003e\u003cp\u003eTwenty-eight patient-oriented questions were collected from Reddit and Quora. A volunteer asked these questions to the five AI chatbots on the same day in a blinded order. Each response was recorded and coded to maintain anonymity. Two prosthodontists independently assessed the responses for accuracy using a 5-point Likert scale, quality using the Global Quality Scale (GQS), and understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT-P). Readability was measured with Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability was calculated using Cohen\u0026rsquo;s kappa. Statistical analyses were performed using Kruskal-Wallis tests for non-parametric data and ANOVA for normally distributed readability scores, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 considered statistically significant.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e\u003cp\u003eSignificant differences were observed in accuracy (p\u0026thinsp;=\u0026thinsp;0.013) and quality (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) among the chatbots. ChatGPT-5 had lower accuracy than Google Gemini (p\u0026thinsp;=\u0026thinsp;0.017) and Claude Sonet 4.5 (p\u0026thinsp;=\u0026thinsp;0.041) and lower quality than all other chatbots (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Readability differed significantly (FRE: p\u0026thinsp;=\u0026thinsp;0.004; FKGL: p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), with ChatGPT-5 responses requiring the highest reading level. PEMAT-P scores also showed significant differences in understandability and actionability (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), with ChatGPT-5 displaying lower scores than other chatbots. Microsoft Copilot, ChatGPT-4, and Google Gemini generally provided higher-quality, more understandable, and actionable information, while ChatGPT-5 and Claude Sonet 4.5 showed limitations. Most chatbot responses were above an eighth-grade reading level, which may challenge general patient comprehension.\u003c/p\u003e\u003ch2\u003eConclusion:\u003c/h2\u003e\u003cp\u003eAI chatbots vary considerably in the quality and usefulness of information they provide for complex dental procedures like smile design. While some models deliver accurate and comprehensible responses, others may produce lower-quality, less actionable content. Despite high understandability in most responses, high reading levels and low actionability could limit patient comprehension and effective decision-making. Care should be taken when patients rely on AI chatbots for dental education, and further improvements are needed to enhance reliability, readability, and actionable guidance.\u003c/p\u003e","manuscriptTitle":"A Comparative Analysis of Five AI Chatbots in Providing Patient Education on Smile Design","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-01 08:43:30","doi":"10.21203/rs.3.rs-8210813/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"235b8c7e-fae1-45c5-add3-aa64cf585611","owner":[],"postedDate":"December 1st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-12-01T08:43:33+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-01 08:43:30","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8210813","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8210813","identity":"rs-8210813","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.