Multimodal Diagnostic Performance of Large Language Models in Dental Image Interpretation: A Comparative Study with Undergraduate Dental Students

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background The rapid integration of artificial intelligence into healthcare and education has intensified interest in the diagnostic capabilities of large language models (LLMs), particularly with the emergence of multimodal systems capable of processing visual data. However, their ability to accurately interpret dental radiographic and clinical images remains uncertain. Methods This comparative observational study evaluated the diagnostic performance of two LLMs (ChatGPT and Google Gemini) against undergraduate dental students. A total of 129 fourth- and fifth-year students assessed 20 diagnostic cases based on panoramic, periapical, cone-beam computed tomography (CBCT), and clinical oral images. AI models were tested using standardized zero-shot prompts with identical inputs over five consecutive days to assess temporal variability. Diagnostic accuracy, response rates, and inter-day consistency were analyzed using chi-square tests, Cochran’s Q test, and Cohen’s kappa coefficient. Results Students demonstrated significantly higher diagnostic accuracy (81.9%) compared with Gemini (63%) and ChatGPT (60%) (p = 0.015). Across all imaging modalities, students consistently outperformed AI models (p < 0.01). ChatGPT achieved a 100% response rate but showed lower accuracy, whereas Gemini demonstrated higher accuracy among answered responses but failed to respond in 21% of cases. AI performance varied significantly across repeated prompts (45%–70%; p = 0.036), with only moderate inter-day agreement (κ = 0.52). Students were 1.78 times more likely to provide correct diagnoses than AI models (p = 0.014). Conclusions Despite recent advances in multimodal capabilities, current LLMs demonstrate limited diagnostic accuracy and inconsistent performance in image-based dental interpretation compared with trained human learners. While these systems may support educational processes, their variability, incomplete response behavior, and reduced diagnostic reliability restrict their use as independent clinical decision-making tools. Further refinement of multimodal AI systems is required before safe integration into dental diagnostics can be considered.
Full text 104,829 characters · extracted from preprint-html · click to expand
Multimodal Diagnostic Performance of Large Language Models in Dental Image Interpretation: A Comparative Study with Undergraduate Dental Students | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Multimodal Diagnostic Performance of Large Language Models in Dental Image Interpretation: A Comparative Study with Undergraduate Dental Students Burcu Yeliz KOLLAYAN, Gizem ALPAY, Mehmet Egemen AYDEMİR This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9193501/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 18 You are reading this latest preprint version Abstract Background The rapid integration of artificial intelligence into healthcare and education has intensified interest in the diagnostic capabilities of large language models (LLMs), particularly with the emergence of multimodal systems capable of processing visual data. However, their ability to accurately interpret dental radiographic and clinical images remains uncertain. Methods This comparative observational study evaluated the diagnostic performance of two LLMs (ChatGPT and Google Gemini) against undergraduate dental students. A total of 129 fourth- and fifth-year students assessed 20 diagnostic cases based on panoramic, periapical, cone-beam computed tomography (CBCT), and clinical oral images. AI models were tested using standardized zero-shot prompts with identical inputs over five consecutive days to assess temporal variability. Diagnostic accuracy, response rates, and inter-day consistency were analyzed using chi-square tests, Cochran’s Q test, and Cohen’s kappa coefficient. Results Students demonstrated significantly higher diagnostic accuracy (81.9%) compared with Gemini (63%) and ChatGPT (60%) (p = 0.015). Across all imaging modalities, students consistently outperformed AI models (p < 0.01). ChatGPT achieved a 100% response rate but showed lower accuracy, whereas Gemini demonstrated higher accuracy among answered responses but failed to respond in 21% of cases. AI performance varied significantly across repeated prompts (45%–70%; p = 0.036), with only moderate inter-day agreement (κ = 0.52). Students were 1.78 times more likely to provide correct diagnoses than AI models (p = 0.014). Conclusions Despite recent advances in multimodal capabilities, current LLMs demonstrate limited diagnostic accuracy and inconsistent performance in image-based dental interpretation compared with trained human learners. While these systems may support educational processes, their variability, incomplete response behavior, and reduced diagnostic reliability restrict their use as independent clinical decision-making tools. Further refinement of multimodal AI systems is required before safe integration into dental diagnostics can be considered. large language models dental education radiographic interpretation artificial intelligence CBCT INTRODUCTION Dental education requires the development of both theoretical knowledge and clinical diagnostic skills. In many countries, including Türkiye, dental education is structured as a five-year program integrating foundational sciences, preclinical training, and clinical practice. Preclinical education typically begins in the third year, whereas clinical training continues throughout the fourth and fifth years. During this period, students receive detailed instruction in dentomaxillofacial anatomy, including radiographic landmarks and the topographic anatomy of the oral mucosa. These courses aim to provide students with the essential knowledge required to recognize anatomical structures and pathological findings in clinical settings. However, students may experience difficulties in recognizing and retaining certain anatomical landmarks, which may complicate the accurate identification of these structures in clinical practice( 1 , 2 ). In recent years, artificial intelligence (AI) technologies have rapidly advanced and have increasingly been integrated into healthcare and medical education( 3 ). AI has demonstrated considerable potential in improving diagnostic accuracy, clinical decision-making, and educational support systems. Among recent AI developments, LLMs such as ChatGPT and Google Gemini have attracted significant attention due to their ability to generate human-like responses and process complex medical information. These models are based on transformer architectures and utilize reinforcement learning from human feedback to improve response quality and contextual understanding ( 4 – 6 ). Recent advances in multimodal artificial intelligence systems have further expanded the capabilities of LLMs. These systems are now able to process not only textual information but also visual data, allowing them to analyze and interpret medical images in addition to generating text-based responses ( 7 ). This development has created new opportunities for applying AI tools in clinical education and diagnostic training. Several studies have investigated the use of LLMs in medical and healthcare education. Previous research has shown that AI-based language models can perform well in answering medical examination questions and may serve as supportive tools in educational settings ( 7 ). However, concerns remain regarding the accuracy, reliability, and consistency of AI-generated responses, particularly in clinical contexts where diagnostic precision is essential. Artificial intelligence applications have also begun to emerge in dentistry, particularly in areas such as radiographic interpretation, diagnostic support systems, and treatment planning ( 3 , 8 – 14 ). Nevertheless, most existing studies in dental education have primarily focused on students’ ability to identify specific anatomical landmarks on panoramic radiographs ( 1 , 2 , 15 ). Studies that evaluate diagnostic interpretation by considering both radiographic images and clinical oral findings simultaneously remain limited. However, the diagnostic performance of multimodal LLMs in image-based dental interpretation, particularly in comparison with dental students, remains insufficiently explored. Understanding the diagnostic capabilities of these systems is particularly important as students increasingly rely on AI tools for learning and clinical guidance. Therefore, the aim of this study was to evaluate the diagnostic performance of large language models in interpreting dental radiographic and clinical images and to compare their performance with that of undergraduate dental students. In addition, the study examined the temporal variability and reliability of AI responses across repeated prompts, providing further insight into the potential role and limitations of LLMs in dental education. The null hypothesis of the study was that different AI chatbot models (ChatGPT and Google Gemini) would respond to diagnostic questions with accuracy comparable to dental students, and that no significant difference would exist between the chatbot models. MATERIALS AND METHODS Study design This study was designed as a comparative observational study aimed at evaluating the diagnostic performance of LLMs in interpreting dental diagnostic images and comparing their performance with that of undergraduate dental students. The study evaluated the interpretation of both dental radiographic images and clinical oral images and the diagnostic accuracy of LLMs was compared with that of dental students. A total of 129 undergraduate dental students participated in the study. Participants were recruited from the fourth and fifth years of the dental program. Demographic and educational information including gender, academic year, and completion of radiology training were recorded for each participant. Participation was voluntary and all responses were collected anonymously. Image dataset and diagnostic questions The study consisted of 20 diagnostic questions based on dental diagnostic images. The images were categorized into four groups: Panoramic radiographs (n = 5) Periapical radiographs (n = 5) Cone-beam computed tomography (CBCT) images (n = 5) Clinical oral images (n = 5) The questionnaire used in this study was specifically developed by the authors and was not adapted from any previously published source. The diagnostic questions were designed based on clinical and radiographic scenarios and were reviewed by two oral and maxillofacial radiologists with more than 5 years of clinical experience (BYK & MEA) to ensure content validity, with a specific focus on anatomical structures and diagnostic scenarios that are frequently misinterpreted by undergraduate dental students, thereby enhancing the educational relevance of the assessment; an English version of the questionnaire is provided as a supplementary file. AI models evaluated Two large language models were evaluated in the study: ChatGPT (Open AI, GPT-5-based model) Google Gemini (Google LLC) Each diagnostic question was submitted to the models using identical prompts. The diagnostic images were directly uploaded to the AI models for evaluation. To evaluate the temporal stability of AI responses, the same set of questions was asked once daily for five consecutive days, resulting in a total of 100 AI responses (20 questions × 5 days). Responses that did not contain a clear answer were categorized as unanswered. All questions were submitted to the AI models using the same standardized prompt format. Each prompt included the diagnostic image and the following instruction: "Please examine the image and select the most appropriate answer among the provided options." The same prompts were used across all testing sessions to ensure consistency and reproducibility of AI responses. Data collection Student responses were collected using a structured questionnaire format. Each participant answered the same set of 20 diagnostic questions. Responses were categorized as correct or incorrect according to the predetermined reference answers. AI model responses were categorized as correct, incorrect, or unanswered. Outcome measures Outcome measures The primary outcome of the study was diagnostic accuracy, defined as the proportion of correct answers relative to the total number of questions. Secondary outcomes included: Response rate of AI models Diagnostic accuracy by imaging modality Comparison of diagnostic accuracy between students and AI models Temporal variability of AI responses Inter-day consistency of AI responses Statistical analysis All statistical analyses were performed using SPSS 23 software ( IBM Corporation, Armonk, NY, USA ) . Descriptive statistics were calculated for participant characteristics and diagnostic accuracy. Differences in diagnostic accuracy between groups were evaluated using the chi-square test. Comparisons between two independent groups (such as academic year and radiology training status) were analyzed using the chi-square test. The temporal variability of AI responses across days was assessed using the Cochran’s Q test. Inter-day consistency of AI responses was evaluated using Cohen’s kappa coefficient. A p-value < 0.05 was considered statistically significant. Ethics statement This study was approved by the Non-Interventional Clinical Research Ethics Committee of Mehmet Akif Ersoy University (Protocol number: GO 2026/2452) and designed in accordance with the Declaration of Helsinki. Participation in the study was voluntary and all data were collected anonymously. RESULTS Participant characteristics A total of 129 dental students participated in the study. Among the participants, 80 (62.0%) were female and 49 (38.0%) were male. Regarding academic level, 76 students (58.9%) were fourth-year students and 53 (41.1%) were fifth-year students. Additionally, 85 students (65.9%) had completed radiology training, while 44 students (34.1%) had not yet received formal radiology training (Table 1 ). Table 1 Participant characteristics of the study population (n = 129) Variable n % Female 80 62.0 Male 49 38.0 4th year students 76 58.9 5th year students 53 41.1 Radiology training completed 85 65.9 Radiology training not completed 44 34.1 Overall student diagnostic performance The overall diagnostic performance of the students demonstrated a mean score of 16.39 ± 2.05 correct answers out of 20 questions, corresponding to a mean diagnostic accuracy of 81.9%. The lowest observed score among participants was 10, whereas the highest score reached 20, indicating variability in diagnostic performance across the cohort (Table 2 ). Table 2 Overall diagnostic performance of dental students Metric Value Number of students 129 Mean correct answers 16.39 ± 2.05 Mean accuracy 81.9% Minimum score 10 Maximum score 20 Question-level diagnostic accuracy Analysis of question-level performance revealed variability in diagnostic accuracy across the 20 questions. The accuracy ranged from 60.5% to 99.2%. The highest accuracy was observed for Question 8 (99.2%), followed by Question 6 (94.6%) and Question 7 (91.5%). In contrast, the lowest accuracy rates were observed for Question 14 (60.5%), Question 10 (61.2%), and Question 18 (65.9%), indicating that these items represented the most challenging diagnostic tasks for the students. Diagnostic accuracy by imaging modality When diagnostic performance was evaluated according to imaging modality, the highest accuracy was observed for periapical radiographs (87.0%), followed by CBCT images (81.0%) and panoramic radiographs (81.1%). The lowest diagnostic accuracy was observed for clinical photographs (77.7%), suggesting that interpretation of clinical lesions may present greater diagnostic challenges for students (Table 3 ). Table 3 Diagnostic accuracy of dental students according to imaging modality Imaging modality Student accuracy (%) Panoramic 81.1 Periapical 87.0 CBCT 81.0 Clinical 77.7 AI model performance Two large language models were evaluated for comparison with student performance. Gemini responded to 79 out of 100 prompts, corresponding to a response rate of 79%. Among the answered prompts, 63 responses were correct, yielding an accuracy of 79.7% among answered questions and an overall accuracy of 63%. In contrast, ChatGPT responded to all prompts (response rate 100%), but demonstrated lower diagnostic accuracy, with 60 correct responses, corresponding to an overall accuracy of 60% (Table 4 ). Table 4 Performance of AI models in diagnostic tasks Model Response rate Accuracy (answered) Overall accuracy Gemini 79% 79.7% 63% ChatGPT 100% 60% 60% * Overall accuracy includes unanswered responses. Comparison between students and AI models Comparative analysis demonstrated that students significantly outperformed AI models in diagnostic accuracy. The overall student accuracy was 81.9%, compared with 63% for Gemini and 60% for ChatGPT. Statistical analysis using the chi-square test revealed a significant difference between students and AI models (p = 0.015) (Table 5 ). Table 5 Comparison of diagnostic accuracy between students and AI models Group Accuracy (%) Students 81.9 Gemini 63 ChatGPT 60 * χ² test, p = 0.015 Diagnostic accuracy by imaging modality: students vs AI Further analysis showed that students consistently achieved higher diagnostic accuracy than AI models across all imaging modalities. Significant differences were observed for panoramic radiographs (χ² = 7.21, p = 0.007), periapical radiographs (χ² = 9.34, p = 0.002), CBCT images (χ² = 10.11, p = 0.001), and clinical photographs (χ² = 8.46, p = 0.004) (Table 6 ). Table 6 Diagnostic accuracy by imaging modality: comparison between students and AI models Imaging modality Students (%) AI (%) χ² p Panoramic 81 64 7.21 0.007 Periapical 87 68 9.34 0.002 CBCT 82 61 10.11 0.001 Clinical 78 56 8.46 0.004 * AI models represent combined performance of ChatGPT and Gemini. Students were significantly more likely to provide correct diagnoses than AI models (OR = 1.78, 95% CI: 1.12–2.83, p = 0.014), indicating a substantial performance advantage. No statistically significant difference in diagnostic accuracy was observed between fourth-year and fifth-year students (p = 0.32). Similarly, completion of radiology training did not significantly affect diagnostic accuracy among students (p = 0.58). Temporal variability of AI performance Evaluation of AI responses across repeated prompts revealed variability in diagnostic performance across different days. Accuracy ranged from 45% to 70%, with the highest performance observed on Day 4 and the lowest on Day 5. Cochran’s Q test demonstrated that this variation was statistically significant (Q = 10.24, p = 0.036) (Table 7 ). Table 7 Temporal variability of AI diagnostic accuracy across repeated prompts Day Accuracy (%) Day 1 60 Day 2 65 Day 3 55 p = 0.036 Day 4 70 Day 5 45 * Cochran’s Q test, p = 0.036 Inter-day reliability analysis demonstrated moderate agreement in AI responses across repeated prompts, with a Cohen’s kappa value of 0.52, indicating moderate consistency in AI-generated diagnostic outputs. DISCUSSION The findings of this study demonstrate a clear and consistent performance gap between undergraduate dental students and current LLMs in the interpretation of dental radiographic and clinical images. Students achieved significantly higher diagnostic accuracy than both evaluated AI models, indicating that current LLMs remain insufficient for complex diagnostic tasks requiring integration of visual information and clinical context. This difference can be explained by the nature of diagnostic reasoning in dentistry. Accurate interpretation of dental images requires not only visual pattern recognition but also clinical reasoning, contextual integration, and decision-making under uncertainty. These competencies are developed through structured education and clinical experience. In contrast, LLMs generate responses based on probabilistic associations derived from large datasets and are primarily optimized for linguistic coherence rather than diagnostic accuracy. As a result, they may fail to capture subtle clinical cues and contextual relationships that are critical for correct diagnosis. Previous studies have reported promising performance of LLMs in medical and dental education, particularly in knowledge-based tasks ( 5 , 9 , 13 , 14 , 22 ). Advanced models such as GPT-4 have demonstrated high accuracy in medical examinations, and ChatGPT has been shown to reach moderate performance levels in standardized tests such as the USMLE ( 5 , 22 ). Similarly, studies in radiology have reported that LLMs can perform well in text-based, exam-style questions ( 16 , 17 ). However, these findings primarily reflect text-based reasoning. The present study indicates that LLM performance is highly task-dependent and decreases substantially in image-based diagnostic tasks. Evidence from dental radiology-specific studies further supports this observation. Öçbe and Demirdurak reported that AI models demonstrated lower accuracy than dental students in identifying anatomical structures on panoramic radiographs ( 1 ). Similarly, Kahalian et al. showed that ChatGPT-4 had limited accuracy in interpreting complex radiographic images ( 18 ). Additional studies have also reported that LLMs tend to generate generalized responses rather than detailed diagnostic interpretations in radiographic analysis ( 23 , 24 ). Together with the present findings, these results indicate that LLMs remain limited in handling image-based diagnostic tasks despite recent advances in multimodal capabilities. Image-based diagnostic interpretation requires spatial reasoning, visual discrimination, and integration of clinical context-processes that differ fundamentally from the text-based mechanisms of LLMs. Although multimodal AI systems have improved visual processing capabilities, their performance in complex diagnostic scenarios remains insufficient when compared with trained human learners. In contrast, task-specific artificial intelligence systems, particularly those based on convolutional neural networks (CNNs), have demonstrated high accuracy in dental imaging tasks such as segmentation and lesion detection ( 19 – 21 ). These findings highlight a critical distinction between general-purpose language models and domain-specific AI systems, suggesting that task-specific models may be more suitable for complex visual diagnostic applications. Another important finding of this study is the variability of AI responses across repeated prompts. Significant fluctuations in diagnostic accuracy were observed across different days, and inter-day agreement was only moderate. This variability raises critical concerns regarding the reproducibility and reliability of AI-generated outputs. In clinical settings, inconsistent responses to identical inputs may lead to diagnostic uncertainty and may pose potential risks to patient safety. From an educational perspective, these findings suggest that LLMs may serve as supportive tools rather than replacements for traditional dental education. While AI systems can facilitate access to information and support learning processes, their limitations in diagnostic accuracy and consistency necessitate cautious use. Furthermore, overreliance on AI-generated outputs may impair the development of clinical reasoning skills and may lead to overconfidence or incorrect diagnostic approaches among students. This study has several limitations. The number of diagnostic images and questions was limited, and only two LLMs were evaluated. Additionally, the use of a standardized prompt may have influenced AI responses. Future studies should include larger datasets, additional AI models, and alternative prompting strategies to further explore the role of LLMs in dental education and clinical diagnostics. Taken together, these findings indicate that although LLMs represent a promising development in medical education, their current limitations in diagnostic accuracy, consistency, and contextual reasoning restrict their use in image-based diagnostic tasks. At present, LLMs should not be considered reliable standalone diagnostic tools in clinical settings. Finally, it emphasizes the importance of maintaining human supervision, particularly in image-based clinical decision-making processes and AI-assisted diagnostic workflows. CONCLUSION In conclusion, large language models demonstrated limited diagnostic accuracy and moderate reliability in interpreting dental radiographic and clinical images compared with undergraduate dental students. Although these systems show potential as supportive educational tools, they currently lack the consistency and contextual understanding required for clinical decision-making. Therefore, LLMs should be integrated cautiously into dental education as complementary resources rather than replacements for clinical training and expert judgment. Future developments in multimodal AI may enhance their diagnostic capabilities and educational value. Declarations Acknowledgements The authors would like to thank all undergraduate dental students who voluntarily participated in this study for their valuable contributions. Author contributions BYK: Conceptualization, Methodology, Investigation, Writing – Original Draft, Writing – Review & Editing. GA: Conceptualization, Methodology, Investigation, Data Curation, Writing - Review & Editing. MEA: Conceptualization, Methodology, Investigation, Writing – Original Draft, Writing – Review & Editing. Funding statement The author(s) received no financial support for the research, authorship, and/or publication of this article. Data availability statement The datasets used and/or analysed during the current study are available from the corresponding author upon reasonable request. Ethics Approval This study was conducted using anonymized patient records and was approved by the Non-Interventional Clinical Research Ethics Committee of Mehmet Akif Ersoy University (Protocol number: GO 2026/2452 ). All procedures were carried out in accordance with the ethical principles of the Declaration of Helsinki. Consent to Participate Informed consent was obtained electronically from all participants prior to participation. The first page of the online questionnaire (Google Forms) included an informed consent statement, and only participants who provided consent were allowed to proceed with the study. Consent for publication Not applicable. Competing interests The author(s) declare no competing interests. Clinical trial number: Not applicable. References Öçbe M, Demirdurak Y. Dentistry Students vs. ChatGPT 4o: Assessing Knowledge of Panoramic Radiographic Anatomy. Aydın Dent J 31 Ağustos. 2025;11(2):139–46. Maeda N, Hosoki H, Yoshida M, Suito H, Honda E. Dental students’ levels of understanding normal panoramic anatomy. J Dent Sci Aralık. 2018;13(4):374–7. 10.1016/j.jds.2018.08 . .002 PubMed PMID: 30895148; PubMed Central PMCID: PMC6388826. Safi Z, Abd-Alrazaq A, Khalifa M, Househ M. Technical Aspects of Developing Chatbots for Medical Applications: Scoping Review. J Med Internet Res 18 Aralık. 2020;22(12):e19127. doi:10.2196/19127 PubMed PMID: 33337337; PubMed Central PMCID: PMC7775817. Han Z, Battaglia F, Udaiyar A, Fooks A, Terlecky SR. An explorative assessment of ChatGPT as an aid in medical education: Use it with caution. Med Teach Mayıs. 2024;46(5):657–64. 2023.2271159 PubMed PMID: 37862566. Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C. vd. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health Şubat. 2023;2(2):e0000198. 10.1371/journal.pdig.0000198 . PubMed PMID: 36812645; PubMed Central PMCID: PMC9931230. Suta P, Lan X, Wu B, Mongkolnam P, Chan JH. An Overview of Machine Learning in Chatbots. Int J Mech Eng Robot Res. 2020;502 – 10. 10.18178/ijmerr.9.4.502-510 Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW. vd. Large language models encode clinical knowledge. Nat Ağustos. 2023;620(7972):172–80. 10.1038/s41586-023-06291-2 . PubMed PMID: 37438534; PubMed Central PMCID: PMC10396962. Eggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent Off Publ Am Acad Esthet Dent Al Ekim. 2023;35(7):1098–102. 10.1111/jerd.13046 . PubMed PMID: 37017291. Sabri H, Saleh MHA, Hazrati P, Merchant K, Misch J, Kumar PS. vd. Performance of three artificial intelligence (AI)-based large language models in standardized testing; implications for AI‐assisted dental education. J Periodontal Res Şubat. 2025;60(2):121–33. 10.1111/jre.13323 . PubMed PMID: 39030766; PubMed Central PMCID: PMC11873669. Fawaz P, Sayegh PE, Vannet BV. What is the current state of artificial intelligence applications in dentistry and orthodontics? J Stomatol Oral Maxillofac Surg Ekim. 2023;124(5):101524. 10.1016/j.jormas.2023.101524 . PubMed PMID: 37270174. Daraqel B, Wafaie K, Mohammed H, Cao L, Mheissen S, Liu Y, vd. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop Off Publ Am Assoc Orthod Its Const Soc Am Board Orthod. Haziran. 2024;165(6):652 – 62. 10.1016/j.ajodo.2024.01 .012 PubMed PMID: 38493370. Makrygiannakis MA, Giannakopoulos K, Kaklamanos EG. Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing. Eur J Orthod 13 Nisan. 2024;48(1):cjae017. 10.1093/ejo/cjae017 . PubMed PMID: 38613510; PubMed Central PMCID: PMC12810200. Revilla-León M, Barmak BA, Sailer I, Kois JC, Att W. Performance of an Artificial Intelligence-Based Chatbot (ChatGPT) Answering the European Certification in Implant Dentistry Exam. Int J Prosthodont 22 Nisan. 2024;37(2):221–4. 10.11607/ijp.8852 . PubMed PMID: 38270461. Bulut AC, Bahadır HS, Ateş G. Artificial intelligence in dental education: can AI-based chatbots compete with general practitioners? BMC Med Educ 02 Ekim. 2025;25:1319. 10.1186/s12909-025-07880-7 . PubMed PMID: 41039365; PubMed Central PMCID: PMC12492586. Razmus TF, Williamson GF, Van Dis ML. Assessment of the knowledge of graduating American dental students about the panoramic image. Oral Surg Oral Med Oral Pathol Eylül. 1993;76(3):397–402. 10.1016/0030-4220 . (93)90278-c PubMed PMID: 8378061. Peker RB. Assessment of the Performance of Large Language Models ChatGPT, ChatGPT Plus, Gemini, and Microsoft Copilot in Responding to Oral and Maxillofacial Radiology Questions from the Turkish Dentistry Specialty Entrance Exam. Yeditepe Dent J. 2025;21(3):130–5. 10.5505/yeditepe.2025.93265 . Bhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations. Radiol Haziran. 2023;307(5):e230582. 10.1148/radiol.230582 . PubMed PMID: 37191485. Kahalian S, Rajabzadeh M, Öçbe M, Medisoglu MS. ChatGPT-4.0 in oral and maxillofacial radiology: prediction of anatomical and pathological conditions from radiographic images. Folia Med (Plovdiv) 31 Aralık. 2024;66(6):863–8. 10.3897/folmed.66.e135584 . Bayrakdar IS, Orhan K, Çelik Ö, Bilgir E, Sağlam H, Kaplan FA. vd. A U-Net Approach to Apical Lesion Segmentation on Panoramic Radiographs. BioMed Res Int. 2022;2022:7035367. 10.1155/2022/7035367 . PubMed PMID: 35075428; PubMed Central PMCID: PMC8783705. Kaygısız Yiğit M, Pınarbaşı A, Etöz M, Duman ŞB, Bayrakdar İŞ. Artificial intelligence-based fully automatic 3D paranasal sinus segmentation. Dento Maxillo Facial Radiol 01 Ocak. 2026;55(1):61–72. 10.1093/dmfr/twaf057 . PubMed PMID: 40711942. Gülşen İT, Kuran A, Evli C, Baydar O, Dinç Başar K, Bilgir E. vd. Deep learning model for automated segmentation of sphenoid sinus and middle skull base structures in CBCT volumes using nnU-Net v2. Oral Radiol Ocak. 2026;42(1):139–47. 10.1007/s11282-025-00848-9 . PubMed PMID: 40748555. Takagi S, Watari T, Erabi A, Sakaguchi K. Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: Comparison Study. JMIR Med Educ 29 Haziran. 2023;9:e48002. doi:10.2196/48002 PubMed PMID: 37384388; PubMed Central PMCID: PMC10365615. Silva TP, Andrade-Bortoletto MFS, Ocampo TSC, Alencar-Palha C, Bornstein MM, Oliveira-Santos C, vd. Performance of a commercially available Generative Pre-trained Transformer (GPT) in describing radiolucent lesions in panoramic radiographs and establishing differential diagnoses. Clin Oral Investig 09 Mart. 2024;28(3):204. 10.1007/s00784-024-05587-5 . PubMed PMID: 38459362; PubMed Central PMCID: PMC10924032. Hu Y, Hu Z, Liu W, Gao A, Wen S, Liu S. vd. Exploring the potential of ChatGPT as an adjunct for generating diagnosis based on chief complaint and cone beam CT radiologic findings. BMC Med Inf Decis Mak 19 Şubat. 2024;24(1):55. 10.1186/s12911-024-02445- . y PubMed PMID: 38374067; PubMed Central PMCID: PMC10875853. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFile.pdf Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 10 May, 2026 Reviews received at journal 10 May, 2026 Reviews received at journal 06 May, 2026 Reviews received at journal 06 May, 2026 Reviewers agreed at journal 04 May, 2026 Reviews received at journal 03 May, 2026 Reviews received at journal 02 May, 2026 Reviewers agreed at journal 30 Apr, 2026 Reviewers agreed at journal 29 Apr, 2026 Reviewers agreed at journal 29 Apr, 2026 Reviewers agreed at journal 29 Apr, 2026 Reviewers agreed at journal 23 Apr, 2026 Reviewers agreed at journal 22 Apr, 2026 Reviewers invited by journal 22 Apr, 2026 Editor assigned by journal 01 Apr, 2026 Editor invited by journal 28 Mar, 2026 Submission checks completed at journal 27 Mar, 2026 First submitted to journal 27 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9193501","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":633597982,"identity":"34a99369-c130-45d1-bd6c-054a81bd0282","order_by":0,"name":"Burcu Yeliz KOLLAYAN","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0klEQVRIiWNgGAWjYBACAygtB2EYWBCvxdiAgRnElSBeS+IGsBYGIrSYSyQfk66ouZO+nb3/6IYfBRIM/O3dCXi1WM5IS5M8c+xZ7s6ew2w3e4AOkzhzdgN+h93IMTZsYDucu+FGMtsNHqAWA4lcQlryPxs2/DucbgDUcvMPcVpyGB82th1OAGm5TZQtlj3PDB829h023HDmsNltGQMJHoJ+MWdPfnCw4dtheYPjjc9uvvljI8ff3otfC4NAAiqfB79yEOA/QFjNKBgFo2AUjHAAAI5SSkLSErLgAAAAAElFTkSuQmCC","orcid":"","institution":"Burdur Mehmet Akif Ersoy University","correspondingAuthor":true,"prefix":"","firstName":"Burcu","middleName":"Yeliz","lastName":"KOLLAYAN","suffix":""},{"id":633597983,"identity":"00974728-ef00-484f-af69-7f76a4caa9dd","order_by":1,"name":"Gizem ALPAY","email":"","orcid":"","institution":"Burdur Mehmet Akif Ersoy University","correspondingAuthor":false,"prefix":"","firstName":"Gizem","middleName":"","lastName":"ALPAY","suffix":""},{"id":633597984,"identity":"b8ebff9b-cdc9-432e-8dd1-96572e729317","order_by":2,"name":"Mehmet Egemen AYDEMİR","email":"","orcid":"","institution":"Burdur Mehmet Akif Ersoy University","correspondingAuthor":false,"prefix":"","firstName":"Mehmet","middleName":"Egemen","lastName":"AYDEMİR","suffix":""}],"badges":[],"createdAt":"2026-03-22 20:08:43","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9193501/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9193501/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109067445,"identity":"185edd9a-42eb-4e43-aec5-99dd196cc3e3","added_by":"auto","created_at":"2026-05-12 09:52:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":286983,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9193501/v1/a066250f-be5a-4e1d-af05-7f2208256b63.pdf"},{"id":108398256,"identity":"ac444e0f-1451-419e-98e1-1b5a70dead6b","added_by":"auto","created_at":"2026-05-04 08:28:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1728991,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFile.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9193501/v1/042f2c7b1f2797f81d26fdfc.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Multimodal Diagnostic Performance of Large Language Models in Dental Image Interpretation: A Comparative Study with Undergraduate Dental Students","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eDental education requires the development of both theoretical knowledge and clinical diagnostic skills. In many countries, including T\u0026uuml;rkiye, dental education is structured as a five-year program integrating foundational sciences, preclinical training, and clinical practice. Preclinical education typically begins in the third year, whereas clinical training continues throughout the fourth and fifth years. During this period, students receive detailed instruction in dentomaxillofacial anatomy, including radiographic landmarks and the topographic anatomy of the oral mucosa. These courses aim to provide students with the essential knowledge required to recognize anatomical structures and pathological findings in clinical settings. However, students may experience difficulties in recognizing and retaining certain anatomical landmarks, which may complicate the accurate identification of these structures in clinical practice(\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn recent years, artificial intelligence (AI) technologies have rapidly advanced and have increasingly been integrated into healthcare and medical education(\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). AI has demonstrated considerable potential in improving diagnostic accuracy, clinical decision-making, and educational support systems. Among recent AI developments, LLMs such as ChatGPT and Google Gemini have attracted significant attention due to their ability to generate human-like responses and process complex medical information. These models are based on transformer architectures and utilize reinforcement learning from human feedback to improve response quality and contextual understanding (\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eRecent advances in multimodal artificial intelligence systems have further expanded the capabilities of LLMs. These systems are now able to process not only textual information but also visual data, allowing them to analyze and interpret medical images in addition to generating text-based responses (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). This development has created new opportunities for applying AI tools in clinical education and diagnostic training.\u003c/p\u003e \u003cp\u003eSeveral studies have investigated the use of LLMs in medical and healthcare education. Previous research has shown that AI-based language models can perform well in answering medical examination questions and may serve as supportive tools in educational settings (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). However, concerns remain regarding the accuracy, reliability, and consistency of AI-generated responses, particularly in clinical contexts where diagnostic precision is essential.\u003c/p\u003e \u003cp\u003eArtificial intelligence applications have also begun to emerge in dentistry, particularly in areas such as radiographic interpretation, diagnostic support systems, and treatment planning (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan additionalcitationids=\"CR9 CR10 CR11 CR12 CR13\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). Nevertheless, most existing studies in dental education have primarily focused on students\u0026rsquo; ability to identify specific anatomical landmarks on panoramic radiographs (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Studies that evaluate diagnostic interpretation by considering both radiographic images and clinical oral findings simultaneously remain limited. However, the diagnostic performance of multimodal LLMs in image-based dental interpretation, particularly in comparison with dental students, remains insufficiently explored.\u003c/p\u003e \u003cp\u003eUnderstanding the diagnostic capabilities of these systems is particularly important as students increasingly rely on AI tools for learning and clinical guidance. Therefore, the aim of this study was to evaluate the diagnostic performance of large language models in interpreting dental radiographic and clinical images and to compare their performance with that of undergraduate dental students. In addition, the study examined the temporal variability and reliability of AI responses across repeated prompts, providing further insight into the potential role and limitations of LLMs in dental education.\u003c/p\u003e \u003cp\u003eThe null hypothesis of the study was that different AI chatbot models (ChatGPT and Google Gemini) would respond to diagnostic questions with accuracy comparable to dental students, and that no significant difference would exist between the chatbot models.\u003c/p\u003e"},{"header":"MATERIALS AND METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy design\u003c/h2\u003e \u003cp\u003eThis study was designed as a comparative observational study aimed at evaluating the diagnostic performance of LLMs in interpreting dental diagnostic images and comparing their performance with that of undergraduate dental students.\u003c/p\u003e \u003cp\u003eThe study evaluated the interpretation of both dental radiographic images and clinical oral images and the diagnostic accuracy of LLMs was compared with that of dental students.\u003c/p\u003e \u003cp\u003eA total of 129 undergraduate dental students participated in the study. Participants were recruited from the fourth and fifth years of the dental program.\u003c/p\u003e \u003cp\u003eDemographic and educational information including gender, academic year, and completion of radiology training were recorded for each participant.\u003c/p\u003e \u003cp\u003eParticipation was voluntary and all responses were collected anonymously.\u003c/p\u003e \u003cp\u003eImage dataset and diagnostic questions\u003c/p\u003e \u003cp\u003eThe study consisted of 20 diagnostic questions based on dental diagnostic images.\u003c/p\u003e \u003cp\u003eThe images were categorized into four groups:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003ePanoramic radiographs (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePeriapical radiographs (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCone-beam computed tomography (CBCT) images (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eClinical oral images (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe questionnaire used in this study was specifically developed by the authors and was not adapted from any previously published source. The diagnostic questions were designed based on clinical and radiographic scenarios and were reviewed by two oral and maxillofacial radiologists with more than 5 years of clinical experience (BYK \u0026amp; MEA) to ensure content validity, with a specific focus on anatomical structures and diagnostic scenarios that are frequently misinterpreted by undergraduate dental students, thereby enhancing the educational relevance of the assessment; an English version of the questionnaire is provided as a supplementary file.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eAI models evaluated\u003c/h3\u003e\n\u003cp\u003eTwo large language models were evaluated in the study:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eChatGPT (Open AI, GPT-5-based model)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eGoogle Gemini (Google LLC)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eEach diagnostic question was submitted to the models using identical prompts. The diagnostic images were directly uploaded to the AI models for evaluation.\u003c/p\u003e \u003cp\u003eTo evaluate the temporal stability of AI responses, the same set of questions was asked once daily for five consecutive days, resulting in a total of 100 AI responses (20 questions \u0026times; 5 days).\u003c/p\u003e \u003cp\u003eResponses that did not contain a clear answer were categorized as unanswered.\u003c/p\u003e \u003cp\u003eAll questions were submitted to the AI models using the same standardized prompt format.\u003c/p\u003e \u003cp\u003eEach prompt included the diagnostic image and the following instruction:\u003c/p\u003e \u003cp\u003e \u003cem\u003e\"Please examine the image and select the most appropriate answer among the provided options.\"\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThe same prompts were used across all testing sessions to ensure consistency and reproducibility of AI responses.\u003c/p\u003e\n\u003ch3\u003eData collection\u003c/h3\u003e\n\u003cp\u003eStudent responses were collected using a structured questionnaire format. Each participant answered the same set of 20 diagnostic questions.\u003c/p\u003e \u003cp\u003eResponses were categorized as correct or incorrect according to the predetermined reference answers.\u003c/p\u003e \u003cp\u003eAI model responses were categorized as correct, incorrect, or unanswered.\u003c/p\u003e\n\u003ch3\u003eOutcome measures\u003c/h3\u003e\n\u003cdiv class=\"Heading\"\u003eOutcome measures\u003c/div\u003e \u003cp\u003eThe primary outcome of the study was diagnostic accuracy, defined as the proportion of correct answers relative to the total number of questions.\u003c/p\u003e \u003cp\u003eSecondary outcomes included:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eResponse rate of AI models\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDiagnostic accuracy by imaging modality\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eComparison of diagnostic accuracy between students and AI models\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTemporal variability of AI responses\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInter-day consistency of AI responses\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses were performed using SPSS 23 software \u003cb\u003e(\u003c/b\u003eIBM Corporation, Armonk, NY, USA\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eDescriptive statistics were calculated for participant characteristics and diagnostic accuracy.\u003c/p\u003e \u003cp\u003eDifferences in diagnostic accuracy between groups were evaluated using the chi-square test.\u003c/p\u003e \u003cp\u003eComparisons between two independent groups (such as academic year and radiology training status) were analyzed using the chi-square test.\u003c/p\u003e \u003cp\u003eThe temporal variability of AI responses across days was assessed using the Cochran\u0026rsquo;s Q test.\u003c/p\u003e \u003cp\u003eInter-day consistency of AI responses was evaluated using Cohen\u0026rsquo;s kappa coefficient.\u003c/p\u003e \u003cp\u003eA p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 was considered statistically significant.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eEthics statement\u003c/h2\u003e \u003cp\u003e This study was approved by the Non-Interventional Clinical Research Ethics Committee of Mehmet Akif Ersoy University (Protocol number: GO 2026/2452) and designed in accordance with the Declaration of Helsinki. Participation in the study was voluntary and all data were collected anonymously.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eParticipant characteristics\u003c/h2\u003e \u003cp\u003eA total of 129 dental students participated in the study. Among the participants, 80 (62.0%) were female and 49 (38.0%) were male. Regarding academic level, 76 students (58.9%) were fourth-year students and 53 (41.1%) were fifth-year students. Additionally, 85 students (65.9%) had completed radiology training, while 44 students (34.1%) had not yet received formal radiology training (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eParticipant characteristics of the study population (n\u0026thinsp;=\u0026thinsp;129)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVariable\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003en\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e%\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFemale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e62.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e38.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4th year students\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e58.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5th year students\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e41.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRadiology training completed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e65.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRadiology training not completed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e34.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eOverall student diagnostic performance\u003c/h2\u003e \u003cp\u003eThe overall diagnostic performance of the students demonstrated a mean score of 16.39\u0026thinsp;\u0026plusmn;\u0026thinsp;2.05 correct answers out of 20 questions, corresponding to a mean diagnostic accuracy of 81.9%. The lowest observed score among participants was 10, whereas the highest score reached 20, indicating variability in diagnostic performance across the cohort (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverall diagnostic performance of dental students\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetric\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValue\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of students\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e129\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMean correct answers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16.39\u0026thinsp;\u0026plusmn;\u0026thinsp;2.05\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMean accuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e81.9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMinimum score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaximum score\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eQuestion-level diagnostic accuracy\u003c/h2\u003e \u003cp\u003eAnalysis of question-level performance revealed variability in diagnostic accuracy across the 20 questions. The accuracy ranged from 60.5% to 99.2%. The highest accuracy was observed for Question 8 (99.2%), followed by Question 6 (94.6%) and Question 7 (91.5%). In contrast, the lowest accuracy rates were observed for Question 14 (60.5%), Question 10 (61.2%), and Question 18 (65.9%), indicating that these items represented the most challenging diagnostic tasks for the students.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eDiagnostic accuracy by imaging modality\u003c/h2\u003e \u003cp\u003eWhen diagnostic performance was evaluated according to imaging modality, the highest accuracy was observed for periapical radiographs (87.0%), followed by CBCT images (81.0%) and panoramic radiographs (81.1%). The lowest diagnostic accuracy was observed for clinical photographs (77.7%), suggesting that interpretation of clinical lesions may present greater diagnostic challenges for students (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDiagnostic accuracy of dental students according to imaging modality\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImaging modality\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStudent accuracy (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePanoramic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e81.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePeriapical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e87.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCBCT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e81.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClinical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e77.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eAI model performance\u003c/h2\u003e \u003cp\u003eTwo large language models were evaluated for comparison with student performance. Gemini responded to 79 out of 100 prompts, corresponding to a response rate of 79%. Among the answered prompts, 63 responses were correct, yielding an accuracy of 79.7% among answered questions and an overall accuracy of 63%.\u003c/p\u003e \u003cp\u003eIn contrast, ChatGPT responded to all prompts (response rate 100%), but demonstrated lower diagnostic accuracy, with 60 correct responses, corresponding to an overall accuracy of 60% (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of AI models in diagnostic tasks\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eResponse rate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy (answered)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOverall accuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e79%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e79.7%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e63%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChatGPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e60%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e60%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003csup\u003e*\u003c/sup\u003eOverall accuracy includes unanswered responses.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eComparison between students and AI models\u003c/h2\u003e \u003cp\u003eComparative analysis demonstrated that students significantly outperformed AI models in diagnostic accuracy. The overall student accuracy was 81.9%, compared with 63% for Gemini and 60% for ChatGPT. Statistical analysis using the chi-square test revealed a significant difference between students and AI models (p\u0026thinsp;=\u0026thinsp;0.015) (Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of diagnostic accuracy between students and AI models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGroup\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStudents\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e81.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGemini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e63\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChatGPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003e\u003csup\u003e*\u003c/sup\u003eχ\u0026sup2; test, p\u0026thinsp;=\u0026thinsp;0.015\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eDiagnostic accuracy by imaging modality: students vs AI\u003c/h2\u003e \u003cp\u003eFurther analysis showed that students consistently achieved higher diagnostic accuracy than AI models across all imaging modalities. Significant differences were observed for panoramic radiographs (χ\u0026sup2; = 7.21, p\u0026thinsp;=\u0026thinsp;0.007), periapical radiographs (χ\u0026sup2; = 9.34, p\u0026thinsp;=\u0026thinsp;0.002), CBCT images (χ\u0026sup2; = 10.11, p\u0026thinsp;=\u0026thinsp;0.001), and clinical photographs (χ\u0026sup2; = 8.46, p\u0026thinsp;=\u0026thinsp;0.004) (Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDiagnostic accuracy by imaging modality: comparison between students and AI models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImaging modality\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStudents (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAI (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eχ\u0026sup2;\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ep\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePanoramic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.007\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePeriapical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.34\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCBCT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClinical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e8.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.004\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"5\"\u003e\u003csup\u003e\u003cb\u003e*\u003c/b\u003e\u003c/sup\u003eAI models represent combined performance of ChatGPT and Gemini.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eStudents were significantly more likely to provide correct diagnoses than AI models (OR\u0026thinsp;=\u0026thinsp;1.78, 95% CI: 1.12\u0026ndash;2.83, p\u0026thinsp;=\u0026thinsp;0.014), indicating a substantial performance advantage.\u003c/p\u003e \u003cp\u003eNo statistically significant difference in diagnostic accuracy was observed between fourth-year and fifth-year students (p\u0026thinsp;=\u0026thinsp;0.32). Similarly, completion of radiology training did not significantly affect diagnostic accuracy among students (p\u0026thinsp;=\u0026thinsp;0.58).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eTemporal variability of AI performance\u003c/h2\u003e \u003cp\u003eEvaluation of AI responses across repeated prompts revealed variability in diagnostic performance across different days. Accuracy ranged from 45% to 70%, with the highest performance observed on Day 4 and the lowest on Day 5. Cochran\u0026rsquo;s Q test demonstrated that this variation was statistically significant (Q\u0026thinsp;=\u0026thinsp;10.24, p\u0026thinsp;=\u0026thinsp;0.036) (Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eTemporal variability of AI diagnostic accuracy across repeated prompts\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay 1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay 2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.036\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay 4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDay 5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"3\"\u003e\u003csup\u003e*\u003c/sup\u003eCochran\u0026rsquo;s Q test, p\u0026thinsp;=\u0026thinsp;0.036\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eInter-day reliability analysis demonstrated moderate agreement in AI responses across repeated prompts, with a Cohen\u0026rsquo;s kappa value of 0.52, indicating moderate consistency in AI-generated diagnostic outputs.\u003c/p\u003e \u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThe findings of this study demonstrate a clear and consistent performance gap between undergraduate dental students and current LLMs in the interpretation of dental radiographic and clinical images. Students achieved significantly higher diagnostic accuracy than both evaluated AI models, indicating that current LLMs remain insufficient for complex diagnostic tasks requiring integration of visual information and clinical context.\u003c/p\u003e \u003cp\u003eThis difference can be explained by the nature of diagnostic reasoning in dentistry. Accurate interpretation of dental images requires not only visual pattern recognition but also clinical reasoning, contextual integration, and decision-making under uncertainty. These competencies are developed through structured education and clinical experience. In contrast, LLMs generate responses based on probabilistic associations derived from large datasets and are primarily optimized for linguistic coherence rather than diagnostic accuracy. As a result, they may fail to capture subtle clinical cues and contextual relationships that are critical for correct diagnosis.\u003c/p\u003e \u003cp\u003ePrevious studies have reported promising performance of LLMs in medical and dental education, particularly in knowledge-based tasks (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). Advanced models such as GPT-4 have demonstrated high accuracy in medical examinations, and ChatGPT has been shown to reach moderate performance levels in standardized tests such as the USMLE (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e). Similarly, studies in radiology have reported that LLMs can perform well in text-based, exam-style questions (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). However, these findings primarily reflect text-based reasoning. The present study indicates that LLM performance is highly task-dependent and decreases substantially in image-based diagnostic tasks.\u003c/p\u003e \u003cp\u003eEvidence from dental radiology-specific studies further supports this observation. \u0026Ouml;\u0026ccedil;be and Demirdurak reported that AI models demonstrated lower accuracy than dental students in identifying anatomical structures on panoramic radiographs (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Similarly, Kahalian et al. showed that ChatGPT-4 had limited accuracy in interpreting complex radiographic images (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e). Additional studies have also reported that LLMs tend to generate generalized responses rather than detailed diagnostic interpretations in radiographic analysis (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e). Together with the present findings, these results indicate that LLMs remain limited in handling image-based diagnostic tasks despite recent advances in multimodal capabilities.\u003c/p\u003e \u003cp\u003eImage-based diagnostic interpretation requires spatial reasoning, visual discrimination, and integration of clinical context-processes that differ fundamentally from the text-based mechanisms of LLMs. Although multimodal AI systems have improved visual processing capabilities, their performance in complex diagnostic scenarios remains insufficient when compared with trained human learners.\u003c/p\u003e \u003cp\u003eIn contrast, task-specific artificial intelligence systems, particularly those based on convolutional neural networks (CNNs), have demonstrated high accuracy in dental imaging tasks such as segmentation and lesion detection (\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). These findings highlight a critical distinction between general-purpose language models and domain-specific AI systems, suggesting that task-specific models may be more suitable for complex visual diagnostic applications.\u003c/p\u003e \u003cp\u003eAnother important finding of this study is the variability of AI responses across repeated prompts. Significant fluctuations in diagnostic accuracy were observed across different days, and inter-day agreement was only moderate. This variability raises critical concerns regarding the reproducibility and reliability of AI-generated outputs. In clinical settings, inconsistent responses to identical inputs may lead to diagnostic uncertainty and may pose potential risks to patient safety.\u003c/p\u003e \u003cp\u003eFrom an educational perspective, these findings suggest that LLMs may serve as supportive tools rather than replacements for traditional dental education. While AI systems can facilitate access to information and support learning processes, their limitations in diagnostic accuracy and consistency necessitate cautious use. Furthermore, overreliance on AI-generated outputs may impair the development of clinical reasoning skills and may lead to overconfidence or incorrect diagnostic approaches among students.\u003c/p\u003e \u003cp\u003eThis study has several limitations. The number of diagnostic images and questions was limited, and only two LLMs were evaluated. Additionally, the use of a standardized prompt may have influenced AI responses. Future studies should include larger datasets, additional AI models, and alternative prompting strategies to further explore the role of LLMs in dental education and clinical diagnostics.\u003c/p\u003e \u003cp\u003eTaken together, these findings indicate that although LLMs represent a promising development in medical education, their current limitations in diagnostic accuracy, consistency, and contextual reasoning restrict their use in image-based diagnostic tasks. At present, LLMs should not be considered reliable standalone diagnostic tools in clinical settings. Finally, it emphasizes the importance of maintaining human supervision, particularly in image-based clinical decision-making processes and AI-assisted diagnostic workflows.\u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eIn conclusion, large language models demonstrated limited diagnostic accuracy and moderate reliability in interpreting dental radiographic and clinical images compared with undergraduate dental students. Although these systems show potential as supportive educational tools, they currently lack the consistency and contextual understanding required for clinical decision-making. Therefore, LLMs should be integrated cautiously into dental education as complementary resources rather than replacements for clinical training and expert judgment. Future developments in multimodal AI may enhance their diagnostic capabilities and educational value.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to thank all undergraduate dental students who voluntarily participated in this study for their valuable contributions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBYK: Conceptualization, Methodology, Investigation, Writing \u0026ndash; Original Draft, Writing \u0026ndash; Review \u0026amp; Editing.\u003c/p\u003e\n\u003cp\u003eGA: Conceptualization, Methodology, Investigation, Data Curation, Writing - Review \u0026amp; Editing.\u003c/p\u003e\n\u003cp\u003eMEA: Conceptualization, Methodology, Investigation, Writing \u0026ndash; Original Draft, Writing \u0026ndash; Review \u0026amp; Editing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author(s) received no financial support for the research, authorship, and/or publication of this article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used and/or analysed during the current study are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics Approval\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was conducted using anonymized patient records and was approved by the Non-Interventional Clinical Research Ethics Committee of Mehmet Akif Ersoy University (Protocol number: \u003cstrong\u003eGO 2026/2452\u003c/strong\u003e). All procedures were carried out in accordance with the ethical principles of the Declaration of Helsinki.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eInformed consent was obtained electronically from all participants prior to participation. The first page of the online questionnaire (Google Forms) included an informed consent statement, and only participants who provided consent were allowed to proceed with the study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author(s) declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003e\u0026Ouml;\u0026ccedil;be M, Demirdurak Y. Dentistry Students vs. ChatGPT 4o: Assessing Knowledge of Panoramic Radiographic Anatomy. Aydın Dent J 31 Ağustos. 2025;11(2):139\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaeda N, Hosoki H, Yoshida M, Suito H, Honda E. Dental students\u0026rsquo; levels of understanding normal panoramic anatomy. J Dent Sci Aralık. 2018;13(4):374\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jds.2018.08\u003c/span\u003e\u003cspan address=\"10.1016/j.jds.2018.08\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. .002 PubMed PMID: 30895148; PubMed Central PMCID: PMC6388826.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSafi Z, Abd-Alrazaq A, Khalifa M, Househ M. Technical Aspects of Developing Chatbots for Medical Applications: Scoping Review. J Med Internet Res 18 Aralık. 2020;22(12):e19127. doi:10.2196/19127 PubMed PMID: 33337337; PubMed Central PMCID: PMC7775817.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan Z, Battaglia F, Udaiyar A, Fooks A, Terlecky SR. An explorative assessment of ChatGPT as an aid in medical education: Use it with caution. Med Teach Mayıs. 2024;46(5):657\u0026ndash;64. 2023.2271159 PubMed PMID: 37862566.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepa\u0026ntilde;o C. vd. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health Şubat. 2023;2(2):e0000198. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pdig.0000198\u003c/span\u003e\u003cspan address=\"10.1371/journal.pdig.0000198\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 36812645; PubMed Central PMCID: PMC9931230.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuta P, Lan X, Wu B, Mongkolnam P, Chan JH. An Overview of Machine Learning in Chatbots. Int J Mech Eng Robot Res. 2020;502\u0026thinsp;\u0026ndash;\u0026thinsp;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.18178/ijmerr.9.4.502-510\u003c/span\u003e\u003cspan address=\"10.18178/ijmerr.9.4.502-510\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW. vd. Large language models encode clinical knowledge. Nat Ağustos. 2023;620(7972):172\u0026ndash;80. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41586-023-06291-2\u003c/span\u003e\u003cspan address=\"10.1038/s41586-023-06291-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 37438534; PubMed Central PMCID: PMC10396962.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent Off Publ Am Acad Esthet Dent Al Ekim. 2023;35(7):1098\u0026ndash;102. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/jerd.13046\u003c/span\u003e\u003cspan address=\"10.1111/jerd.13046\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 37017291.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSabri H, Saleh MHA, Hazrati P, Merchant K, Misch J, Kumar PS. vd. Performance of three artificial intelligence (AI)-based large language models in standardized testing; implications for AI‐assisted dental education. J Periodontal Res Şubat. 2025;60(2):121\u0026ndash;33. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/jre.13323\u003c/span\u003e\u003cspan address=\"10.1111/jre.13323\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 39030766; PubMed Central PMCID: PMC11873669.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFawaz P, Sayegh PE, Vannet BV. What is the current state of artificial intelligence applications in dentistry and orthodontics? J Stomatol Oral Maxillofac Surg Ekim. 2023;124(5):101524. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jormas.2023.101524\u003c/span\u003e\u003cspan address=\"10.1016/j.jormas.2023.101524\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 37270174.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDaraqel B, Wafaie K, Mohammed H, Cao L, Mheissen S, Liu Y, vd. The performance of artificial intelligence models in generating responses to general orthodontic questions: ChatGPT vs Google Bard. Am J Orthod Dentofac Orthop Off Publ Am Assoc Orthod Its Const Soc Am Board Orthod. Haziran. 2024;165(6):652\u0026thinsp;\u0026ndash;\u0026thinsp;62. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ajodo.2024.01\u003c/span\u003e\u003cspan address=\"10.1016/j.ajodo.2024.01\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.012 PubMed PMID: 38493370.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMakrygiannakis MA, Giannakopoulos K, Kaklamanos EG. Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing. Eur J Orthod 13 Nisan. 2024;48(1):cjae017. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/ejo/cjae017\u003c/span\u003e\u003cspan address=\"10.1093/ejo/cjae017\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 38613510; PubMed Central PMCID: PMC12810200.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRevilla-Le\u0026oacute;n M, Barmak BA, Sailer I, Kois JC, Att W. Performance of an Artificial Intelligence-Based Chatbot (ChatGPT) Answering the European Certification in Implant Dentistry Exam. Int J Prosthodont 22 Nisan. 2024;37(2):221\u0026ndash;4. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.11607/ijp.8852\u003c/span\u003e\u003cspan address=\"10.11607/ijp.8852\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 38270461.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBulut AC, Bahadır HS, Ateş G. Artificial intelligence in dental education: can AI-based chatbots compete with general practitioners? BMC Med Educ 02 Ekim. 2025;25:1319. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12909-025-07880-7\u003c/span\u003e\u003cspan address=\"10.1186/s12909-025-07880-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 41039365; PubMed Central PMCID: PMC12492586.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRazmus TF, Williamson GF, Van Dis ML. Assessment of the knowledge of graduating American dental students about the panoramic image. Oral Surg Oral Med Oral Pathol Eyl\u0026uuml;l. 1993;76(3):397\u0026ndash;402. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/0030-4220\u003c/span\u003e\u003cspan address=\"10.1016/0030-4220\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. (93)90278-c PubMed PMID: 8378061.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeker RB. Assessment of the Performance of Large Language Models ChatGPT, ChatGPT Plus, Gemini, and Microsoft Copilot in Responding to Oral and Maxillofacial Radiology Questions from the Turkish Dentistry Specialty Entrance Exam. Yeditepe Dent J. 2025;21(3):130\u0026ndash;5. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5505/yeditepe.2025.93265\u003c/span\u003e\u003cspan address=\"10.5505/yeditepe.2025.93265\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhayana R, Krishna S, Bleakney RR. Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations. Radiol Haziran. 2023;307(5):e230582. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.230582\u003c/span\u003e\u003cspan address=\"10.1148/radiol.230582\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 37191485.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKahalian S, Rajabzadeh M, \u0026Ouml;\u0026ccedil;be M, Medisoglu MS. ChatGPT-4.0 in oral and maxillofacial radiology: prediction of anatomical and pathological conditions from radiographic images. Folia Med (Plovdiv) 31 Aralık. 2024;66(6):863\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3897/folmed.66.e135584\u003c/span\u003e\u003cspan address=\"10.3897/folmed.66.e135584\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBayrakdar IS, Orhan K, \u0026Ccedil;elik \u0026Ouml;, Bilgir E, Sağlam H, Kaplan FA. vd. A U-Net Approach to Apical Lesion Segmentation on Panoramic Radiographs. BioMed Res Int. 2022;2022:7035367. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1155/2022/7035367\u003c/span\u003e\u003cspan address=\"10.1155/2022/7035367\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 35075428; PubMed Central PMCID: PMC8783705.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaygısız Yiğit M, Pınarbaşı A, Et\u0026ouml;z M, Duman ŞB, Bayrakdar İŞ. Artificial intelligence-based fully automatic 3D paranasal sinus segmentation. Dento Maxillo Facial Radiol 01 Ocak. 2026;55(1):61\u0026ndash;72. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/dmfr/twaf057\u003c/span\u003e\u003cspan address=\"10.1093/dmfr/twaf057\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 40711942.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG\u0026uuml;lşen İT, Kuran A, Evli C, Baydar O, Din\u0026ccedil; Başar K, Bilgir E. vd. Deep learning model for automated segmentation of sphenoid sinus and middle skull base structures in CBCT volumes using nnU-Net v2. Oral Radiol Ocak. 2026;42(1):139\u0026ndash;47. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11282-025-00848-9\u003c/span\u003e\u003cspan address=\"10.1007/s11282-025-00848-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 40748555.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakagi S, Watari T, Erabi A, Sakaguchi K. Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: Comparison Study. JMIR Med Educ 29 Haziran. 2023;9:e48002. doi:10.2196/48002 PubMed PMID: 37384388; PubMed Central PMCID: PMC10365615.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSilva TP, Andrade-Bortoletto MFS, Ocampo TSC, Alencar-Palha C, Bornstein MM, Oliveira-Santos C, vd. Performance of a commercially available Generative Pre-trained Transformer (GPT) in describing radiolucent lesions in panoramic radiographs and establishing differential diagnoses. Clin Oral Investig 09 Mart. 2024;28(3):204. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00784-024-05587-5\u003c/span\u003e\u003cspan address=\"10.1007/s00784-024-05587-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 38459362; PubMed Central PMCID: PMC10924032.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu Y, Hu Z, Liu W, Gao A, Wen S, Liu S. vd. Exploring the potential of ChatGPT as an adjunct for generating diagnosis based on chief complaint and cone beam CT radiologic findings. BMC Med Inf Decis Mak 19 Şubat. 2024;24(1):55. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12911-024-02445-\u003c/span\u003e\u003cspan address=\"10.1186/s12911-024-02445-\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. y PubMed PMID: 38374067; PubMed Central PMCID: PMC10875853.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"large language models, dental education, radiographic interpretation, artificial intelligence, CBCT","lastPublishedDoi":"10.21203/rs.3.rs-9193501/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9193501/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe rapid integration of artificial intelligence into healthcare and education has intensified interest in the diagnostic capabilities of large language models (LLMs), particularly with the emergence of multimodal systems capable of processing visual data. However, their ability to accurately interpret dental radiographic and clinical images remains uncertain.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThis comparative observational study evaluated the diagnostic performance of two LLMs (ChatGPT and Google Gemini) against undergraduate dental students. A total of 129 fourth- and fifth-year students assessed 20 diagnostic cases based on panoramic, periapical, cone-beam computed tomography (CBCT), and clinical oral images. AI models were tested using standardized zero-shot prompts with identical inputs over five consecutive days to assess temporal variability. Diagnostic accuracy, response rates, and inter-day consistency were analyzed using chi-square tests, Cochran\u0026rsquo;s Q test, and Cohen\u0026rsquo;s kappa coefficient.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eStudents demonstrated significantly higher diagnostic accuracy (81.9%) compared with Gemini (63%) and ChatGPT (60%) (p\u0026thinsp;=\u0026thinsp;0.015). Across all imaging modalities, students consistently outperformed AI models (p\u0026thinsp;\u0026lt;\u0026thinsp;0.01). ChatGPT achieved a 100% response rate but showed lower accuracy, whereas Gemini demonstrated higher accuracy among answered responses but failed to respond in 21% of cases. AI performance varied significantly across repeated prompts (45%\u0026ndash;70%; p\u0026thinsp;=\u0026thinsp;0.036), with only moderate inter-day agreement (κ\u0026thinsp;=\u0026thinsp;0.52). Students were 1.78 times more likely to provide correct diagnoses than AI models (p\u0026thinsp;=\u0026thinsp;0.014).\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eDespite recent advances in multimodal capabilities, current LLMs demonstrate limited diagnostic accuracy and inconsistent performance in image-based dental interpretation compared with trained human learners. While these systems may support educational processes, their variability, incomplete response behavior, and reduced diagnostic reliability restrict their use as independent clinical decision-making tools. Further refinement of multimodal AI systems is required before safe integration into dental diagnostics can be considered.\u003c/p\u003e","manuscriptTitle":"Multimodal Diagnostic Performance of Large Language Models in Dental Image Interpretation: A Comparative Study with Undergraduate Dental Students","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-04 08:27:56","doi":"10.21203/rs.3.rs-9193501/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-10T12:11:39+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-10T08:15:04+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-06T07:15:47+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-06T06:46:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"293608274657054087918323970693440766609","date":"2026-05-04T05:24:30+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-03T11:28:55+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-02T17:52:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"252871235012440321136429867398928533108","date":"2026-04-30T13:15:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"271057353682467283981726555874779705183","date":"2026-04-29T09:40:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"159497876095653928667323459691649481175","date":"2026-04-29T07:01:52+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"45528107166407350137487930075162312866","date":"2026-04-29T07:01:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"9182063199003533369717916470147946312","date":"2026-04-23T12:44:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"139130280449630423475328810081830120535","date":"2026-04-23T03:40:54+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-22T18:32:19+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-01T05:58:54+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-28T20:03:46+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-27T11:47:12+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2026-03-27T11:41:58+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"56f96c76-5124-468c-8ff8-146516723ebc","owner":[],"postedDate":"May 4th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-10T12:11:39+00:00","index":71,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-10T08:15:04+00:00","index":70,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-06T07:15:47+00:00","index":69,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-06T06:46:12+00:00","index":68,"fulltext":""},{"type":"reviewerAgreed","content":"293608274657054087918323970693440766609","date":"2026-05-04T05:24:30+00:00","index":66,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-03T11:28:55+00:00","index":65,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-02T17:52:50+00:00","index":64,"fulltext":""},{"type":"reviewerAgreed","content":"252871235012440321136429867398928533108","date":"2026-04-30T13:15:23+00:00","index":62,"fulltext":""},{"type":"reviewerAgreed","content":"271057353682467283981726555874779705183","date":"2026-04-29T09:40:27+00:00","index":61,"fulltext":""},{"type":"reviewerAgreed","content":"159497876095653928667323459691649481175","date":"2026-04-29T07:01:52+00:00","index":59,"fulltext":""},{"type":"reviewerAgreed","content":"45528107166407350137487930075162312866","date":"2026-04-29T07:01:39+00:00","index":58,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-04T08:27:56+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-04 08:27:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9193501","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9193501","identity":"rs-9193501","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00