Evaluating Large Language Models for Translating Caries Guidelines into Clinical Decision Support

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Objective To systematically evaluate the capability of three large language models (LLMs)—ChatGPT-4o, Grok-3, and DeepSeek—in interpreting and translating clinical practice guidelines for caries management and in supporting clinical decision-making, thereby exploring their potential role in disseminating dental knowledge and assisting clinical practice. Methods Based on the American Dental Association’s Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions, a zero-shot prompting strategy was used to instruct each model to generate guideline summaries tailored for both healthcare professionals and the general public. Additionally, the models were asked to provide diagnoses and treatment plans for three standardized clinical cases related to caries. Manual evaluations were conducted across five dimensions—accuracy, clarity, conciseness, logical coherence, and overall quality—using a 0–10 scoring system. Text consistency was also assessed using ROUGE-L and BLEU metrics. Results In generating guideline summaries for the general public, GPT-4o achieved the highest overall score, excelling particularly in clarity and logical coherence, while DeepSeek performed best in terminology accuracy and fidelity to the source text. For summaries intended for healthcare professionals, all three models performed well, with DeepSeek leading in automated evaluation metrics. In clinical case management, ChatGPT attained the highest composite score, significantly outperforming Grok-3 and DeepSeek, demonstrating superior diagnostic accuracy and clinically relevant treatment recommendations. Conclusion Large language models show promising potential in translating dental guidelines and assisting clinical decision-making. However, limitations such as insufficient personalization and mechanistic application of guidelines remain. Future efforts should focus on integrating multimodal data, enabling dynamic knowledge updates, and developing human–AI collaborative care models to achieve a balance between standardized and personalized management of oral diseases.
Full text 70,848 characters · extracted from preprint-html · click to expand
Evaluating Large Language Models for Translating Caries Guidelines into Clinical Decision Support | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluating Large Language Models for Translating Caries Guidelines into Clinical Decision Support Gu Nan, Bingxin Fan, Yao Yuan, Xinliang Duan, Sichen Han, Zhenyong Tang, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8542862/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 27 Apr, 2026 Read the published version in BMC Oral Health → Version 1 posted 10 You are reading this latest preprint version Abstract Objective To systematically evaluate the capability of three large language models (LLMs)—ChatGPT-4o, Grok-3, and DeepSeek—in interpreting and translating clinical practice guidelines for caries management and in supporting clinical decision-making, thereby exploring their potential role in disseminating dental knowledge and assisting clinical practice. Methods Based on the American Dental Association’s Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions, a zero-shot prompting strategy was used to instruct each model to generate guideline summaries tailored for both healthcare professionals and the general public. Additionally, the models were asked to provide diagnoses and treatment plans for three standardized clinical cases related to caries. Manual evaluations were conducted across five dimensions—accuracy, clarity, conciseness, logical coherence, and overall quality—using a 0–10 scoring system. Text consistency was also assessed using ROUGE-L and BLEU metrics. Results In generating guideline summaries for the general public, GPT-4o achieved the highest overall score, excelling particularly in clarity and logical coherence, while DeepSeek performed best in terminology accuracy and fidelity to the source text. For summaries intended for healthcare professionals, all three models performed well, with DeepSeek leading in automated evaluation metrics. In clinical case management, ChatGPT attained the highest composite score, significantly outperforming Grok-3 and DeepSeek, demonstrating superior diagnostic accuracy and clinically relevant treatment recommendations. Conclusion Large language models show promising potential in translating dental guidelines and assisting clinical decision-making. However, limitations such as insufficient personalization and mechanistic application of guidelines remain. Future efforts should focus on integrating multimodal data, enabling dynamic knowledge updates, and developing human–AI collaborative care models to achieve a balance between standardized and personalized management of oral diseases. Large Language Models Clinical Decision Support Guideline Adherence Caries Management Figures Figure 1 Figure 2 Figure 3 Introduction Dental caries is a highly prevalent global disease and a primary initiating factor for pulp and periapical pathologies. The standardized management of caries is therefore crucial for preventing the progression to more complex endodontic conditions [ 1 – 3 ] . Pulp and periapical diseases themselves are common occurrences in dentistry. The accuracy of their diagnosis is closely associated with treatment efficacy, directly impacting patients' masticatory function, oral health, and quality of life [ 4 – 6 ] . For a long time, the diagnosis of these conditions has been heavily reliant on clinicians' experience and the subjective interpretation of radiographic findings, leading to variability among practitioners of different experience levels and affecting the consistency and accuracy of diagnosis and treatment [ 7 – 12 ] . The treatment process also faces challenges such as the complex anatomy of the root canal system and the limitations of equipment precision. Furthermore, postoperative prognosis assessment is often based on generalized experience, lacking personalized predictive tools. In recent years, the rapid advancement of digital technology has brought about a revolutionary shift in the diagnosis of endodontic diseases [ 13 – 15 ] . Three-dimensional imaging technologies, such as intraoral scanning, have significantly improved the visualization of tooth structure. Meanwhile, artificial intelligence, particularly the widespread application of deep learning models in medical image analysis, has provided novel approaches for the automated detection of caries and periapical lesions and for assisted decision-making, substantially enhancing diagnostic efficiency and standardization [ 16 ] . Within this context, Large Language Models (LLMs), as a pivotal branch of natural language processing, have demonstrated the potential to integrate and translate specialized medical knowledge, promising to further advance the precision and accessibility of oral healthcare [ 17 ] . Based on the American Dental Association’s Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions and the Australian Society of Endodontology's Guidelines for Non-Surgical Root Canal Treatment, this study selected three major AI platforms—GPT-4o (ChatGPT-4o), Grok-3, and DeepSeek—to systematically evaluate their capabilities in interpreting professional guidelines and generating summaries for different target audiences. It also examines their reasoning and decision-support abilities in clinical case management, aiming to explore the application prospects of LLMs in enhancing knowledge dissemination and clinical practice in dentistry [ 18 , 19 ] . Methods 1. Study Design and Guideline Selection This study employed a comparative analytical design to evaluate the capabilities of three advanced LLMs—GPT-4o, Grok-3, and DeepSeek—in interpreting and translating complex clinical practice guidelines into actionable summaries for different target audiences, as well as generating diagnoses and treatment plans for standardized clinical cases. The source documents for this evaluation were the Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions (2018) published by the American Dental Association (ADA) and the Guidelines for Non-Surgical Root Canal Treatment (2024) issued by the Australian Society of Endodontology Inc. These guidelines were selected due to their clinical relevance, timeliness, and structured presentation of evidence-based recommendations, providing a robust benchmark for assessing the LLMs’ knowledge comprehension and translation capabilities within the field of endodontics. 2. Text Translation and Clinical Case Analysis via Large Language Models Between June 1 and June 10, 2025, the three aforementioned LLMs were accessed via their public application programming interfaces (APIs). All interactions were conducted in Chinese to ensure consistency and practical applicability within the primary research context. A standardized zero-shot prompting strategy was applied across all tasks to objectively evaluate the models’ intrinsic capabilities without extensive task-specific fine-tuning. Each model was required to complete two main tasks: 2.1 Guideline Summary Generation Each model was prompted to generate two distinct summaries of the guidelines: one tailored for oral healthcare professionals (e.g., junior dentists, dental students), requiring technical accuracy and inclusion of clinical details; and another aimed at the general public, emphasizing clarity, avoidance of jargon, and practical information. 2.2 Simulated Clinical Case Analysis Following the summary task, each model was provided with three standardized clinical vignettes (brief case descriptions) depicting patients with various complexities of carious lesions and symptoms of pulpitis. The models were instructed to assume the role of an endodontic consultant and provide a differential diagnosis, a step-by-step treatment plan based on guideline recommendations, and detailed postoperative instructions. 3. Evaluation Metrics The outputs of each model were evaluated across five dimensions—accuracy, clarity, conciseness, logical coherence, and overall quality—using a manual scoring system ranging from 0 to 10 points. Additionally, automated metrics including ROUGE-L and BLEU were employed to assess textual consistency. Results 1. Analysis of Medical Documentation (For Non-Medical Audiences) In the task of translating caries management guidelines for non-medical audiences, GPT-4o demonstrated superior performance, achieving an overall score of 9.03 (Table 1). It excelled particularly in clarity and logical coherence. Its content was well-structured and employed effective terminology simplification strategies, such as comparing early caries lesions to “white spot lesions,” which significantly improved public comprehension. DeepSeek scored 8.83 in accuracy and 9.0 in logical coherence (Table 1). Its outputs showed high fidelity to the American Dental Association guidelines, accurately reflecting professional details such as fluoride concentration, the ICDAS classification system, and indications for silver diamine fluoride. However, its use of excessive professional terminology (e.g., remineralization and resin infiltration techniques) posed cognitive barriers for non-expert readers. Although Grok-3 offers advantages in open access and cost-effectiveness, it received a composite score of only 8.37. It scored merely 8.0 and 8.17 in conciseness and clarity, respectively, exhibiting significant issues with content redundancy—such as repeatedly emphasizing sugar reduction—and failed to clearly distinguish treatment priorities between primary and permanent teeth. Automated metric validation revealed that DeepSeek achieved the highest ROUGE-L and BLEU scores (0.89 and 0.85, respectively), indicating the strongest consistency with the original guideline (Table 2, Figure 1e, 1f, 2d, 2e, 2f). While GPT-4o performed excellently in linguistic adaptability, it occasionally lacked detail in presenting specific treatment frequencies and sources of evidence. Figure 1a-c and Figure 2a-c visually compare the models' performance profiles across these key dimensions. Overall, current LLMs show significant potential in translating medical guidelines for public use but require further optimization in balancing terminology control, detail completeness, and audience appropriateness. A comprehensive summary of strengths and limitations for each model in this task is provided in Table 5. 2. Analysis of Medical Documentation (For Junior Doctors and Medical Students) In the evaluation of caries guideline translation capabilities for junior dentists and medical students, the outputs of three LLMs—GPT-4o, DeepSeek, and Grok-3—were systematically analyzed across five dimensions: accuracy, clarity, conciseness, logical coherence, and overall quality. The results indicated that all three models achieved high accuracy scores of 9.5 (Table 3), demonstrating strong adherence to the original guideline’s professional recommendations and evidence grades, particularly in rigorously presenting non-restorative treatment recommendations. In terms of clarity, GPT-4o led with a perfect score, offering clear structure and standardized terminology that significantly enhanced comprehension efficiency for medical trainees; both DeepSeek and Grok-3 also performed excellently, scoring 9.5. For conciseness, DeepSeek and Grok-3 scored 9.0, slightly higher than GPT-4o’s 8.5 (Table 3), indicating better information condensation capabilities. In logical coherence, GPT-4o again ranked first with a perfect score by strictly adhering to the clinical reasoning pathway of “etiology – clinical presentation – diagnosis – treatment – prognosis.” DeepSeek followed closely with a score of 9.75, demonstrating excellent reasoning continuity (Table 3). All three models received an overall quality score of 9.5, indicating general reliability in medical knowledge translation. Automated metrics showed that DeepSeek achieved slightly higher ROUGE-L and BLEU scores (0.85 and 0.81, respectively) than GPT-4o and Grok-3 (Table 4), reflecting better semantic and linguistic alignment with the guideline text. In summary, although GPT-4o held a slight advantage in clarity and logical coherence, the two open-source models demonstrated significant potential in comprehensive performance and accessibility, making them particularly suitable for generating professional content in medical education that balances accuracy, comprehensibility, and efficiency. Their distinct profiles are further detailed in Table 5. 3. Analysis of Diagnosis and Treatment in Clinical Cases In the analysis of diagnosis and treatment planning for actual clinical cases, systematic comparison of the three large language models across multiple caries-related cases revealed significant differences in their ability to translate caries management guidelines into clinical practice. Overall, GPT-4o performed optimally in standardized processing protocols, diagnostic accuracy, and rationality of treatment plans, achieving a composite score of 92.51, significantly higher than Grok-3’s 68.67 and DeepSeek’s 50.67 (Table 6, Figure 3e, 3f). Specifically, GPT-4o effectively identified key clinical manifestations and radiographic features in case analyses and provided stratified treatment recommendations based on guidelines. Particularly in complex cases—such as deep caries with pulp exposure or periapical pathology—its proposed root canal treatment and subsequent restorative plans adhered to clinical standards. Although Grok-3 demonstrated some clinical logic in certain cases (e.g., understanding the application of non-restorative treatments like fluoride), its diagnostic details and plan completeness were weaker. It failed to adequately integrate dynamic monitoring and intervention strategies from the guidelines, especially in cases complicated by pulpal pathology. DeepSeek exhibited strong capabilities in using technical terminology, accurately referencing professional concepts such as ICDAS classification and fluoride agent concentrations. However, its recommendations often lacked clinical applicability and demonstrated a mechanistic application of guidelines. This was particularly evident in its failure to correctly distinguish the applicability boundaries between "non-restorative treatments for caries" and "treatments for periapical diseases," leading to knowledge transfer errors. For instance, it mechanically suggested non-restorative treatments for periapical cases without obvious carious lesions, reflecting a limited ability to tailor guideline recommendations to specific case characteristics. The variation in performance across different clinical cases is illustrated in Figure 3c and 3d. Furthermore, in completeness of instructions, GPT-4o also led with a score of 9.17 (Table 6). Its postoperative guidance covered pain management, oral hygiene maintenance, and long-term follow-up requirements. In contrast, Grok-3 and DeepSeek scored only 7.33 and 6.17, respectively, showing significant deficiencies in personalized preventive measures and key patient communication points. The comparative performance across the four core clinical dimensions (standardization, diagnosis, treatment selection, and order completeness) is visually summarized in Figure 3a and 3b. In conclusion, current LLMs still face multiple challenges in clinical reasoning for caries and endodontics, including accurate interpretation of guideline recommendations, generation of individualized treatment strategies, and supporting shared decision-making with patients. Their practical application requires integration with professional judgment to ensure diagnostic and therapeutic quality and patient safety. Conclusions and Future Perspectives This systematic evaluation of three major language models in translating caries and endodontic management guidelines and supporting clinical decision-making demonstrates that current artificial intelligence (AI) technologies remain in an auxiliary stage within oral healthcare practice. Although GPT-4o exhibited notable clinical applicability and outperformed Grok-3 and DeepSeek in standardizing diagnostic and therapeutic workflows, diagnostic accuracy, and treatment rationale, significant limitations persist. The models showed insufficient ability to integrate guideline recommendations with individualized patient management in complex clinical scenarios. This was particularly evident in atypical cases, such as asymptomatic periapical lesions, where DeepSeek mechanically applied non-restorative treatment advice, revealing a lack of deep understanding of pathological mechanisms and an inability to correctly delineate the scope of different guidelines. Future development should focus on constructing multimodal data integration and dynamic knowledge-updating mechanisms. Firstly, efforts should promote the synergistic analysis of LLMs with imaging, molecular diagnostics, and real-time monitoring data. For example, incorporating cone-beam CT (CBCT) imaging features and biomarkers of caries activity into decision-making systems could enhance the accuracy of pulp status interpretation. Current studies have already shown high consistency (exceeding 90% accuracy) between AI and clinicians in detecting and segmenting chronic periapical lesions [ 20 ] . Secondly, it is essential to establish automated pipelines linking guideline updates to model retraining, ensuring treatment recommendations remain synchronized with advances in evidence-based medicine. Simultaneously, greater emphasis must be placed on the complementary roles of clinical experience and AI-driven decision support, fostering the development of human-AI collaborative care models with clinicians at the helm. This will enhance the models' capacity to predict individual variations in treatment outcomes. Furthermore, challenges related to technological generalizability and equity must be addressed to ensure that AI-driven improvements in diagnosis and treatment are accessible even in resource-limited settings. The ultimate goal is to develop intelligent clinical systems that embody precision, adaptability, and ethical compliance, thereby unifying standardized and personalized approaches to the prevention and management of diseases such as dental caries and periapical pathologies. Declarations Conflicts of Interest: The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Funding This research was funded by Natural Science Fund of China, grant number 82301026; Innovation and entrepreneurship program for college students, grant number S202510183727, X202510183467; Postgraduate Plan, grant number 2025CX325; Postgraduate Education and Teaching Reform Research, grant number 2025RGZNZ010; Special Project on AI-Empowered Undergraduate Education and Teaching Reform, grant number 24AI109Z; Jilin Provincial Department of Finance, grant number jcsz2025678-14. Author Contribution N.G. and B.F. contributed equally to this work. Conceptualization: N.G., B.F., Z.W., and J.S. Methodology: N.G., B.F., Y.Y., and X.D. Formal analysis and investigation: S.H., Z.T., and X.D. Writing – original draft preparation: N.G. and B.F. Writing – review and editing: Y.Y., Z.W., and J.S. Supervision and project administration: J.S. and Z.W. All authors read and approved the final manuscript. Data Availability Data will be made available on request. References Siqueira JF Jr, Rôças IN. Present status and future directions: Microbiology of endodontic infections. Int Endod J. 2022;55(Suppl 3):512–30. Wei Y, Lyu P, Bi R, Chen X, Yu Y, Li Z, Fan Y. Neural Regeneration in Regenerative Endodontic Treatment: An Overview and Current Trends. Int J Mol Sci. 2022;23(24):15492. Huang D, Wang X, Liang J, Ling J, Bian Z, Yu Q, Hou B, Chen X, Li J, Ye L, Cheng L, Xu X, Hu T, Wu H, Guo B, Su Q, Chen Z, Qiu L, Chen W, Wei X, Huang Z, Yu J, Lin Z, Zhang Q, Yang D, Zhao J, Pan S, Yang J, Wu J, Pan Y, Xie X, Deng S, Huang X, Zhang L, Yue L, Zhou X. Expert consensus on difficulty assessment of endodontic therapy. Int J Oral Sci. 2024;16(1):22. Yapp KE, Suleiman M, Brennan P, Ekpo E. Periapical Radiography versus Cone Beam Computed Tomography in Endodontic Disease Detection: A Free-response, Factorial Study. J Endod. 2023;49(4):419–29. European Society of Endodontology (ESE) developed by:, Duncan HF, Galler KM, Tomson PL, Simon S, El-Karim I, Kundzina R, Krastl G, Dammaschke T, Fransson H, Markvart M, Zehnder M, Bjørndal L. European Society of Endodontology position statement: Management of deep caries and the exposed pulp. Int Endod J. 2019;52(7):923–34. Estrela C, Decurcio DA, Rossi-Fedele G, Silva JA, Guedes OA, Borges ÁH. Root perforations: a review of diagnosis, prognosis and materials. Braz Oral Res. 2018;32(suppl 1):e73. Kakka A, Gavriil D, Whitworth J. Treatment of cracked teeth: A comprehensive narrative review. Clin Exp Dent Res. 2022;8(5):1218–48. Agrawal P, Nikhade P. Artificial Intelligence in Dentistry: Past, Present, and Future. Cureus. 2022;14(7):e27405. Gulabivala K, Ng YL. Factors that affect the outcomes of root canal treatment and retreatment-A reframing of the principles. Int Endod J. 2023;56(Suppl 2):82–115. Wei X, Du Y, Zhou X, Yue L, Yu Q, Hou B, Chen Z, Liang J, Chen W, Qiu L, Huang X, Meng L, Huang D, Wang X, Tian Y, Tang Z, Zhang Q, Miao L, Zhao J, Yang D, Yang J, Ling J. Expert consensus on digital guided therapy for endodontic diseases. Int J Oral Sci. 2023;15(1):54. Khan R, Akbar S, Khan A, Marwan M, Qaisar ZH, Mehmood A, Shahid F, Munir K, Zheng Z. Dental image enhancement network for early diagnosis of oral dental disease. Sci Rep. 2023;13(1):5312. Xu X, Zheng X, Lin F, Yu Q, Hou B, Chen Z, Wei X, Qiu L, Wenxia C, Li J, Chen L, Wang Z, Wu H, Lu Z, Zhao J, Liang Y, Zhao J, Pan Y, Pan S, Wang X, Yang D, Ren Y, Yue L, Zhou X. Expert consensus on endodontic therapy for patients with systemic conditions. Int J Oral Sci. 2024;16(1):45. Mohammad-Rahimi H, Motamedian SR, Rohban MH, Krois J, Uribe SE, Mahmoudinia E, Rokhshad R, Nadimi M, Schwendicke F. Deep learning for caries detection: A systematic review. J Dent. 2022;122:104115. Sadr S, Mohammad-Rahimi H, Motamedian SR, Zahedrozegar S, Motie P, Vinayahalingam S, Dianat O, Nosrat A. Deep Learning for Detection of Periapical Radiolucent Lesions: A Systematic Review and Meta-analysis of Diagnostic Test Accuracy. J Endod. 2023;49(3):248–e2613. De Rosa CS, Bergamini ML, Palmieri M, Sarmento DJS, de Carvalho MO, Ricardo ALF, Hasseus B, Jonasson P, Braz-Silva PH, Ferreira Costa AL. Differentiation of periapical granuloma from radicular cyst using cone beam computed tomography images texture analysis. Heliyon. 2020;6(10):e05194. Künzle P, Paris S. Performance of large language artificial intelligence models on solving restorative dentistry and endodontics student assessments. Clin Oral Investig. 2024;28(11):575. Huang H, Zheng O, Wang D, Yin J, Wang Z, Ding S, Yin H, Xu C, Yang R, Zheng Q, Shi B. ChatGPT for shaping the future of dentistry: the potential of multi-modal large language model. Int J Oral Sci. 2023;15(1):29. Slayton RL, Urquhart O, Araujo MWB, Fontana M, Guzmán-Armstrong S, Nascimento MM, Nový BB, Tinanoff N, Weyant RJ, Wolff MS, Young DA, Zero DT, Tampi MP, Pilcher L, Banfield L, Carrasco-Labra A. Evidence-based clinical practice guideline on nonrestorative treatments for carious lesions: A report from the American Dental Association. J Am Dent Assoc. 2018;149(10):837–e84919. Peters OA, Rossi-Fedele G, George R, Kumar K, Timmerman A, Wright PP. Guidelines for non-surgical root canal treatment. Aust Endod J. 2024;50(2):202–14. Boubaris M, Cameron A, Manakil J, George R. Artificial intelligence vs. semi-automated segmentation for assessment of dental periapical lesion volume index score: A cone-beam CT study. Comput Biol Med. 2024;175:108527. Tables Table 1 to 6 are available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Table1.docx SupplementalInformation.docx Graphicalabstract.png Cite Share Download PDF Status: Published Journal Publication published 27 Apr, 2026 Read the published version in BMC Oral Health → Version 1 posted Editorial decision: Revision requested 10 Feb, 2026 Reviews received at journal 10 Feb, 2026 Reviews received at journal 27 Jan, 2026 Reviewers agreed at journal 26 Jan, 2026 Reviewers agreed at journal 22 Jan, 2026 Reviewers invited by journal 22 Jan, 2026 Editor invited by journal 22 Jan, 2026 Editor assigned by journal 12 Jan, 2026 Submission checks completed at journal 12 Jan, 2026 First submitted to journal 07 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8542862","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":580569029,"identity":"bf638d0b-3173-42db-bef2-5145f43d27bf","order_by":0,"name":"Gu Nan","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Gu","middleName":"","lastName":"Nan","suffix":""},{"id":580569030,"identity":"f6ec2ec4-d7b3-49e0-8c71-4d8691457980","order_by":1,"name":"Bingxin Fan","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Bingxin","middleName":"","lastName":"Fan","suffix":""},{"id":580569031,"identity":"fa45ba9f-082a-4a83-94aa-ad3031c3c26d","order_by":2,"name":"Yao Yuan","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Yao","middleName":"","lastName":"Yuan","suffix":""},{"id":580569032,"identity":"f19e8168-4c97-404c-bd73-f43cccb5970d","order_by":3,"name":"Xinliang Duan","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Xinliang","middleName":"","lastName":"Duan","suffix":""},{"id":580569033,"identity":"9882165d-f695-436d-9196-5460d536fab9","order_by":4,"name":"Sichen Han","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Sichen","middleName":"","lastName":"Han","suffix":""},{"id":580569034,"identity":"195bda40-27fe-4ace-b0ed-e329047d8f98","order_by":5,"name":"Zhenyong Tang","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Zhenyong","middleName":"","lastName":"Tang","suffix":""},{"id":580569035,"identity":"22a52f9d-1322-42a8-8e2a-ee16ba5b4206","order_by":6,"name":"Jiayu Shen","email":"","orcid":"","institution":"Jilin University","correspondingAuthor":false,"prefix":"","firstName":"Jiayu","middleName":"","lastName":"Shen","suffix":""},{"id":580569036,"identity":"10883998-89a6-4d07-9200-8478b9274485","order_by":7,"name":"Zilin Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA30lEQVRIie3QMQrCMBSA4YhQlxwggt7hQSFVEL1KitCpiKNjwV1n0UP0CE86dKm4VlwyFYQOunUQNFFc07oJ5ifwMrwPQgix2X42GFDSiXxU11bUkDBKKIpvCGHqCNKMQHpILtWc9byulAklo36M7UIaSTYLBkw9bLgVQpHAjdHxwEQ4hhxAETi/SOLHSB1mJMeSg9DkhJo8GpA8dCVqkhNNsJ5M8pK3Ik0yIfY7mLqbxOFG0l2H7q26jyaQZr4sF+P+Kl0WRqL6PIMK9YFqtmv29cr1PTtYv2uz2Wx/2RNkMEqCbXtREwAAAABJRU5ErkJggg==","orcid":"","institution":"Jilin University","correspondingAuthor":true,"prefix":"","firstName":"Zilin","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2026-01-07 14:53:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8542862/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8542862/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12903-026-08430-3","type":"published","date":"2026-04-27T15:58:34+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":101273631,"identity":"2cf6f83e-edb9-4c33-8a69-4a09c64b04ed","added_by":"auto","created_at":"2026-01-28 03:06:52","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1065296,"visible":true,"origin":"","legend":"\u003cp\u003eComparative Performance Evaluation of AI Platforms\u003cstrong\u003e:\u003c/strong\u003e (a) Comparative Evaluation of AI Platforms across Key Dimensions\u003cstrong\u003e;\u003c/strong\u003e(b) 3 Radar Chart of Performance Metrics with Error Bars.\u003cstrong\u003e Caption:\u003c/strong\u003e Comparison of five performance metrics (Accuracy, Clarity, Conciseness, Logic, and Overall Assessment) across three AI platforms. Error bars represent one standard deviation from the mean. The radar chart visualizes the strengths and weaknesses of each platform across different dimensions; (c) Heatmap of Performance Metrics with Significance Annotation.\u003cstrong\u003e Caption:\u003c/strong\u003e Heatmap displaying mean scores for each metric and platform. Numerical values indicate mean scores, and symbols denote statistical significance from simulated ANOVA and post-hoc Tukey HSD tests (*p \u0026lt; 0.05, **p \u0026lt; 0.01, ns: not significant)\u003cstrong\u003e;\u003c/strong\u003e (d) Rose wreath diagram;(e) \u003cstrong\u003eDumbbell Chart Comparing ROUGE-L and BLEU Scores. \u003c/strong\u003eCaption: Dumbbell chart highlighting the difference between ROUGE-L (left points) and BLEU (right points) scores for each platform. The connecting lines visualize the gap between performance on these two related but distinct text evaluation metrics; (f) \u003cstrong\u003eCorrelation Scatter Plot with Regression Line. \u003c/strong\u003eCaption: Scatter plot examining the relationship between ROUGE-L and BLEU scores across platforms. The dashed regression line suggests a positive correlation between these two text evaluation metrics. Platform labels are positioned to avoid overlap using the repel algorithm.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/e7681048da89c4c05ab6de05.jpeg"},{"id":101273628,"identity":"c9297ade-f714-42b5-8c8e-792df268d268","added_by":"auto","created_at":"2026-01-28 03:06:52","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1214993,"visible":true,"origin":"","legend":"\u003cp\u003eComprehensive Multidimensional Performance Evaluation of AI Platforms. (a) \u003cstrong\u003eRadar Chart of Performance Metrics.\u003c/strong\u003e Caption: Radar chart illustrating the performance profile of each platform across evaluation metrics. The polygon area represents overall capability, with DeepSeek showing the most balanced and comprehensive performance across both text evaluation metrics; (b) Violin Plots of Performance Metrics with Statistical Comparisons; (c)\u003cstrong\u003e Parallel Coordinates Plot.\u003c/strong\u003e Caption: Line plot connecting performance across metrics for each platform, allowing visual tracing of performance profiles. ChatGPT-4o shows the most consistent performance across metrics, while DeepSeek demonstrates the highest overall scores; (d) \u003cstrong\u003eComparative Analysis of AI Platform Performance. \u003c/strong\u003eCaption: Comprehensive evaluation of three AI platforms (ChatGPT-4o, Groks, and DeepSeek) using ROUGE-L and BLEU metrics for text generation quality assessment. Multiple visualization techniques provide complementary perspectives on platform performance, highlighting DeepSeek's superior performance across both evaluation metrics; (e) \u003cstrong\u003eDumbbell Chart Comparing Metric Differences Caption.\u003c/strong\u003e Visualization highlighting the difference between ROUGE-L (left points) and BLEU (right points) scores for each platform. The connecting lines illustrate performance gaps between metrics, with all platforms showing higher ROUGE-L than BLEU scores; (f) \u003cstrong\u003eBubble Plot of ROUGE-L vs BLEU with Composite Scoring.\u003c/strong\u003e Caption: Scatter plot comparing ROUGE-L and BLEU scores, with bubble size representing composite performance (average of both metrics). Platforms positioned toward the top-right corner with larger bubbles demonstrate superior overall performance in text generation tasks.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/b9327596487b4685025d1c36.jpeg"},{"id":101297616,"identity":"b9f0df67-cd68-4d4c-8a6a-b4737cc2bc99","added_by":"auto","created_at":"2026-01-28 09:28:18","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":178297,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparative Analysis of Clinical Capabilities Across AI Platforms. (a) Radar Chart of Average Dimension Scores by AI Platform.\u003c/strong\u003e Caption: A radar chart comparing the average scores of AI platforms across four dimensions (Standardization of Treatment, Diagnostic Accuracy, Treatment Plan Selection, and Completeness of Medical Orders). ChatGPT leads in all dimensions, while Grok3 and Deepseek show significant gaps in the \"Completeness of Medical Orders\" dimension; \u003cstrong\u003e(b) Distribution of Dimension Scores by AI Platform.\u003c/strong\u003e Caption: Violin plots display the score distributions of AI platforms across the four dimensions. ChatGPT shows a concentrated and high distribution, while Grok3 and Deepseek exhibit more dispersed distributions with lower tails, particularly in the \"Completeness of Medical Orders\" dimension (right violin); \u003cstrong\u003e(c) Trend of AI Platform Scores Across Clinical Cases.\u003c/strong\u003e Caption: A line chart illustrates the total score variations of AI platforms across different clinical cases. Grok3 performs moderately in most cases; \u003cstrong\u003e(d) Stacked Dimension Scores by Case and AI Platform.\u003c/strong\u003e Caption: Stacked bar charts show the contribution of each dimension to the total score for every case and AI platform combination;\u003cstrong\u003e (e) Comparison of Total Scores Across AI Platforms.\u003c/strong\u003eCaption: Box plots illustrates the distribution of total scores for ChatGPT, Grok3, and Deepseek across six clinical cases. The box represents the interquartile range (IQR), the midline denotes the median, and the dots indicate raw data points. ChatGPT demonstrates the highest score stability, while Deepseek shows greater fluctuations; \u003cstrong\u003e(f) Radar Chart of Average Case Analysis Scores by AI Platform. \u003c/strong\u003eCaption: A radar chart comparing the average scores of AI platforms across dimensions including Standardization of Treatment Process, Clinical Diagnostic Accuracy, Treatment Plan Selection, Completeness of Medical Orders, and Total Score.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/09f353f096215d32e668e170.png"},{"id":108438049,"identity":"a51091ef-96f1-4e34-b3b4-10e19746ec8f","added_by":"auto","created_at":"2026-05-04 16:06:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2612135,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/430296dd-cd9a-43a7-8b33-7b322a924427.pdf"},{"id":101751221,"identity":"2fe97972-c41c-41ff-be7f-9b0c16a34ed2","added_by":"auto","created_at":"2026-02-03 10:18:25","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":22175,"visible":true,"origin":"","legend":"","description":"","filename":"Table1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/b61fc9b9c0f9d6a028ecd784.docx"},{"id":101298117,"identity":"67e40456-6a5a-4a06-b094-59f52f7f4cc0","added_by":"auto","created_at":"2026-01-28 09:30:27","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":176586,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementalInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/b1f94788e88cf66e232b6476.docx"},{"id":101273632,"identity":"adba63f1-620a-4f42-a80a-1827db813168","added_by":"auto","created_at":"2026-01-28 03:06:52","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":75797,"visible":true,"origin":"","legend":"","description":"","filename":"Graphicalabstract.png","url":"https://assets-eu.researchsquare.com/files/rs-8542862/v1/13be1e1c501d6111e4e29a30.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluating Large Language Models for Translating Caries Guidelines into Clinical Decision Support","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDental caries is a highly prevalent global disease and a primary initiating factor for pulp and periapical pathologies. The standardized management of caries is therefore crucial for preventing the progression to more complex endodontic conditions\u003csup\u003e[\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e–\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u003c/sup\u003e. Pulp and periapical diseases themselves are common occurrences in dentistry. The accuracy of their diagnosis is closely associated with treatment efficacy, directly impacting patients' masticatory function, oral health, and quality of life\u003csup\u003e[\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e–\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]\u003c/sup\u003e. For a long time, the diagnosis of these conditions has been heavily reliant on clinicians' experience and the subjective interpretation of radiographic findings, leading to variability among practitioners of different experience levels and affecting the consistency and accuracy of diagnosis and treatment\u003csup\u003e[\u003cspan additionalcitationids=\"CR8 CR9 CR10 CR11\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e–\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]\u003c/sup\u003e. The treatment process also faces challenges such as the complex anatomy of the root canal system and the limitations of equipment precision. Furthermore, postoperative prognosis assessment is often based on generalized experience, lacking personalized predictive tools.\u003c/p\u003e \u003cp\u003eIn recent years, the rapid advancement of digital technology has brought about a revolutionary shift in the diagnosis of endodontic diseases\u003csup\u003e[\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e–\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/sup\u003e. Three-dimensional imaging technologies, such as intraoral scanning, have significantly improved the visualization of tooth structure. Meanwhile, artificial intelligence, particularly the widespread application of deep learning models in medical image analysis, has provided novel approaches for the automated detection of caries and periapical lesions and for assisted decision-making, substantially enhancing diagnostic efficiency and standardization\u003csup\u003e[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]\u003c/sup\u003e. Within this context, Large Language Models (LLMs), as a pivotal branch of natural language processing, have demonstrated the potential to integrate and translate specialized medical knowledge, promising to further advance the precision and accessibility of oral healthcare \u003csup\u003e[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eBased on the American Dental Association’s Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions and the Australian Society of Endodontology's Guidelines for Non-Surgical Root Canal Treatment, this study selected three major AI platforms—GPT-4o (ChatGPT-4o), Grok-3, and DeepSeek—to systematically evaluate their capabilities in interpreting professional guidelines and generating summaries for different target audiences. It also examines their reasoning and decision-support abilities in clinical case management, aiming to explore the application prospects of LLMs in enhancing knowledge dissemination and clinical practice in dentistry\u003csup\u003e[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e "},{"header":"Methods","content":"\u003ch3\u003e1. Study Design and Guideline Selection\u003c/h3\u003e\u003cp\u003eThis study employed a comparative analytical design to evaluate the capabilities of three advanced LLMs—GPT-4o, Grok-3, and DeepSeek—in interpreting and translating complex clinical practice guidelines into actionable summaries for different target audiences, as well as generating diagnoses and treatment plans for standardized clinical cases. The source documents for this evaluation were the Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions (2018) published by the American Dental Association (ADA) and the Guidelines for Non-Surgical Root Canal Treatment (2024) issued by the Australian Society of Endodontology Inc. These guidelines were selected due to their clinical relevance, timeliness, and structured presentation of evidence-based recommendations, providing a robust benchmark for assessing the LLMs’ knowledge comprehension and translation capabilities within the field of endodontics.\u003c/p\u003e\u003ch3\u003e2. Text Translation and Clinical Case Analysis via Large Language Models\u003c/h3\u003e\u003cp\u003eBetween June 1 and June 10, 2025, the three aforementioned LLMs were accessed via their public application programming interfaces (APIs). All interactions were conducted in Chinese to ensure consistency and practical applicability within the primary research context. A standardized zero-shot prompting strategy was applied across all tasks to objectively evaluate the models’ intrinsic capabilities without extensive task-specific fine-tuning. Each model was required to complete two main tasks:\u003c/p\u003e\u003ch2\u003e2.1 Guideline Summary Generation\u003c/h2\u003e\u003cp\u003eEach model was prompted to generate two distinct summaries of the guidelines: one tailored for oral healthcare professionals (e.g., junior dentists, dental students), requiring technical accuracy and inclusion of clinical details; and another aimed at the general public, emphasizing clarity, avoidance of jargon, and practical information.\u003c/p\u003e\u003ch2\u003e2.2 Simulated Clinical Case Analysis\u003c/h2\u003e\u003cp\u003eFollowing the summary task, each model was provided with three standardized clinical vignettes (brief case descriptions) depicting patients with various complexities of carious lesions and symptoms of pulpitis. The models were instructed to assume the role of an endodontic consultant and provide a differential diagnosis, a step-by-step treatment plan based on guideline recommendations, and detailed postoperative instructions.\u003c/p\u003e\u003ch3\u003e3. Evaluation Metrics\u003c/h3\u003e\u003cp\u003eThe outputs of each model were evaluated across five dimensions—accuracy, clarity, conciseness, logical coherence, and overall quality—using a manual scoring system ranging from 0 to 10 points. Additionally, automated metrics including ROUGE-L and BLEU were employed to assess textual consistency.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003e1. Analysis of Medical Documentation (For Non-Medical Audiences)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the task of translating caries management guidelines for non-medical audiences, GPT-4o demonstrated superior performance, achieving an overall score of 9.03 (Table 1). It excelled particularly in clarity and logical coherence. Its content was well-structured and employed effective terminology simplification strategies, such as comparing early caries lesions to \u0026ldquo;white spot lesions,\u0026rdquo; which significantly improved public comprehension. DeepSeek scored 8.83 in accuracy and 9.0 in logical coherence (Table 1). Its outputs showed high fidelity to the American Dental Association guidelines, accurately reflecting professional details such as fluoride concentration, the ICDAS classification system, and indications for silver diamine fluoride. However, its use of excessive professional terminology (e.g., remineralization and resin infiltration techniques) posed cognitive barriers for non-expert readers. Although Grok-3 offers advantages in open access and cost-effectiveness, it received a composite score of only 8.37. It scored merely 8.0 and 8.17 in conciseness and clarity, respectively, exhibiting significant issues with content redundancy\u0026mdash;such as repeatedly emphasizing sugar reduction\u0026mdash;and failed to clearly distinguish treatment priorities between primary and permanent teeth. Automated metric validation revealed that DeepSeek achieved the highest ROUGE-L and BLEU scores (0.89 and 0.85, respectively), indicating the strongest consistency with the original guideline (Table 2, Figure 1e, 1f, 2d, 2e, 2f). While GPT-4o performed excellently in linguistic adaptability, it occasionally lacked detail in presenting specific treatment frequencies and sources of evidence. Figure 1a-c and Figure 2a-c visually compare the models\u0026apos; performance profiles across these key dimensions. Overall, current LLMs show significant potential in translating medical guidelines for public use but require further optimization in balancing terminology control, detail completeness, and audience appropriateness. A comprehensive summary of strengths and limitations for each model in this task is provided in Table 5.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2. Analysis of Medical Documentation (For Junior Doctors and Medical Students)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the evaluation of caries guideline translation capabilities for junior dentists and medical students, the outputs of three LLMs\u0026mdash;GPT-4o, DeepSeek, and Grok-3\u0026mdash;were systematically analyzed across five dimensions: accuracy, clarity, conciseness, logical coherence, and overall quality. The results indicated that all three models achieved high accuracy scores of 9.5 (Table 3), demonstrating strong adherence to the original guideline\u0026rsquo;s professional recommendations and evidence grades, particularly in rigorously presenting non-restorative treatment recommendations. In terms of clarity, GPT-4o led with a perfect score, offering clear structure and standardized terminology that significantly enhanced comprehension efficiency for medical trainees; both DeepSeek and Grok-3 also performed excellently, scoring 9.5. For conciseness, DeepSeek and Grok-3 scored 9.0, slightly higher than GPT-4o\u0026rsquo;s 8.5 (Table 3), indicating better information condensation capabilities. In logical coherence, GPT-4o again ranked first with a perfect score by strictly adhering to the clinical reasoning pathway of \u0026ldquo;etiology \u0026ndash; clinical presentation \u0026ndash; diagnosis \u0026ndash; treatment \u0026ndash; prognosis.\u0026rdquo; DeepSeek followed closely with a score of 9.75, demonstrating excellent reasoning continuity (Table 3). All three models received an overall quality score of 9.5, indicating general reliability in medical knowledge translation. Automated metrics showed that DeepSeek achieved slightly higher ROUGE-L and BLEU scores (0.85 and 0.81, respectively) than GPT-4o and Grok-3 (Table 4), reflecting better semantic and linguistic alignment with the guideline text. In summary, although GPT-4o held a slight advantage in clarity and logical coherence, the two open-source models demonstrated significant potential in comprehensive performance and accessibility, making them particularly suitable for generating professional content in medical education that balances accuracy, comprehensibility, and efficiency. Their distinct profiles are further detailed in Table 5.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3. Analysis of Diagnosis and Treatment in Clinical Cases\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the analysis of diagnosis and treatment planning for actual clinical cases, systematic comparison of the three large language models across multiple caries-related cases revealed significant differences in their ability to translate caries management guidelines into clinical practice. Overall, GPT-4o performed optimally in standardized processing protocols, diagnostic accuracy, and rationality of treatment plans, achieving a composite score of 92.51, significantly higher than Grok-3\u0026rsquo;s 68.67 and DeepSeek\u0026rsquo;s 50.67 (Table 6, Figure 3e, 3f). Specifically, GPT-4o effectively identified key clinical manifestations and radiographic features in case analyses and provided stratified treatment recommendations based on guidelines. Particularly in complex cases\u0026mdash;such as deep caries with pulp exposure or periapical pathology\u0026mdash;its proposed root canal treatment and subsequent restorative plans adhered to clinical standards. Although Grok-3 demonstrated some clinical logic in certain cases (e.g., understanding the application of non-restorative treatments like fluoride), its diagnostic details and plan completeness were weaker. It failed to adequately integrate dynamic monitoring and intervention strategies from the guidelines, especially in cases complicated by pulpal pathology. DeepSeek exhibited strong capabilities in using technical terminology, accurately referencing professional concepts such as ICDAS classification and fluoride agent concentrations. However, its recommendations often lacked clinical applicability and demonstrated a mechanistic application of guidelines. This was particularly evident in its failure to correctly distinguish the applicability boundaries between \u0026quot;non-restorative treatments for caries\u0026quot; and \u0026quot;treatments for periapical diseases,\u0026quot; leading to knowledge transfer errors. For instance, it mechanically suggested non-restorative treatments for periapical cases without obvious carious lesions, reflecting a limited ability to tailor guideline recommendations to specific case characteristics. The variation in performance across different clinical cases is illustrated in Figure 3c and 3d.\u003c/p\u003e\n\u003cp\u003eFurthermore, in completeness of instructions, GPT-4o also led with a score of 9.17 (Table 6). Its postoperative guidance covered pain management, oral hygiene maintenance, and long-term follow-up requirements. In contrast, Grok-3 and DeepSeek scored only 7.33 and 6.17, respectively, showing significant deficiencies in personalized preventive measures and key patient communication points. The comparative performance across the four core clinical dimensions (standardization, diagnosis, treatment selection, and order completeness) is visually summarized in Figure 3a and 3b. In conclusion, current LLMs still face multiple challenges in clinical reasoning for caries and endodontics, including accurate interpretation of guideline recommendations, generation of individualized treatment strategies, and supporting shared decision-making with patients. Their practical application requires integration with professional judgment to ensure diagnostic and therapeutic quality and patient safety.\u003c/p\u003e"},{"header":"Conclusions and Future Perspectives","content":"\u003cp\u003eThis systematic evaluation of three major language models in translating caries and endodontic management guidelines and supporting clinical decision-making demonstrates that current artificial intelligence (AI) technologies remain in an auxiliary stage within oral healthcare practice. Although GPT-4o exhibited notable clinical applicability and outperformed Grok-3 and DeepSeek in standardizing diagnostic and therapeutic workflows, diagnostic accuracy, and treatment rationale, significant limitations persist. The models showed insufficient ability to integrate guideline recommendations with individualized patient management in complex clinical scenarios. This was particularly evident in atypical cases, such as asymptomatic periapical lesions, where DeepSeek mechanically applied non-restorative treatment advice, revealing a lack of deep understanding of pathological mechanisms and an inability to correctly delineate the scope of different guidelines.\u003c/p\u003e\u003cp\u003eFuture development should focus on constructing multimodal data integration and dynamic knowledge-updating mechanisms. Firstly, efforts should promote the synergistic analysis of LLMs with imaging, molecular diagnostics, and real-time monitoring data. For example, incorporating cone-beam CT (CBCT) imaging features and biomarkers of caries activity into decision-making systems could enhance the accuracy of pulp status interpretation. Current studies have already shown high consistency (exceeding 90% accuracy) between AI and clinicians in detecting and segmenting chronic periapical lesions\u003csup\u003e[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]\u003c/sup\u003e. Secondly, it is essential to establish automated pipelines linking guideline updates to model retraining, ensuring treatment recommendations remain synchronized with advances in evidence-based medicine. Simultaneously, greater emphasis must be placed on the complementary roles of clinical experience and AI-driven decision support, fostering the development of human-AI collaborative care models with clinicians at the helm. This will enhance the models' capacity to predict individual variations in treatment outcomes.\u003c/p\u003e\u003cp\u003eFurthermore, challenges related to technological generalizability and equity must be addressed to ensure that AI-driven improvements in diagnosis and treatment are accessible even in resource-limited settings. The ultimate goal is to develop intelligent clinical systems that embody precision, adaptability, and ethical compliance, thereby unifying standardized and personalized approaches to the prevention and management of diseases such as dental caries and periapical pathologies.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eConflicts of Interest:\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis research was funded by Natural Science Fund of China, grant number 82301026; Innovation and entrepreneurship program for college students, grant number S202510183727, X202510183467; Postgraduate Plan, grant number 2025CX325; Postgraduate Education and Teaching Reform Research, grant number 2025RGZNZ010; Special Project on AI-Empowered Undergraduate Education and Teaching Reform, grant number 24AI109Z; Jilin Provincial Department of Finance, grant number jcsz2025678-14.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eN.G. and B.F. contributed equally to this work. Conceptualization: N.G., B.F., Z.W., and J.S. Methodology: N.G., B.F., Y.Y., and X.D. Formal analysis and investigation: S.H., Z.T., and X.D. Writing \u0026ndash; original draft preparation: N.G. and B.F. Writing \u0026ndash; review and editing: Y.Y., Z.W., and J.S. Supervision and project administration: J.S. and Z.W. All authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eData will be made available on request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSiqueira JF Jr, R\u0026ocirc;\u0026ccedil;as IN. Present status and future directions: Microbiology of endodontic infections. Int Endod J. 2022;55(Suppl 3):512\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei Y, Lyu P, Bi R, Chen X, Yu Y, Li Z, Fan Y. Neural Regeneration in Regenerative Endodontic Treatment: An Overview and Current Trends. Int J Mol Sci. 2022;23(24):15492.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang D, Wang X, Liang J, Ling J, Bian Z, Yu Q, Hou B, Chen X, Li J, Ye L, Cheng L, Xu X, Hu T, Wu H, Guo B, Su Q, Chen Z, Qiu L, Chen W, Wei X, Huang Z, Yu J, Lin Z, Zhang Q, Yang D, Zhao J, Pan S, Yang J, Wu J, Pan Y, Xie X, Deng S, Huang X, Zhang L, Yue L, Zhou X. Expert consensus on difficulty assessment of endodontic therapy. Int J Oral Sci. 2024;16(1):22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYapp KE, Suleiman M, Brennan P, Ekpo E. Periapical Radiography versus Cone Beam Computed Tomography in Endodontic Disease Detection: A Free-response, Factorial Study. J Endod. 2023;49(4):419\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEuropean Society of Endodontology (ESE) developed by:, Duncan HF, Galler KM, Tomson PL, Simon S, El-Karim I, Kundzina R, Krastl G, Dammaschke T, Fransson H, Markvart M, Zehnder M, Bj\u0026oslash;rndal L. European Society of Endodontology position statement: Management of deep caries and the exposed pulp. Int Endod J. 2019;52(7):923\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEstrela C, Decurcio DA, Rossi-Fedele G, Silva JA, Guedes OA, Borges \u0026Aacute;H. Root perforations: a review of diagnosis, prognosis and materials. Braz Oral Res. 2018;32(suppl 1):e73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKakka A, Gavriil D, Whitworth J. Treatment of cracked teeth: A comprehensive narrative review. Clin Exp Dent Res. 2022;8(5):1218\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgrawal P, Nikhade P. Artificial Intelligence in Dentistry: Past, Present, and Future. Cureus. 2022;14(7):e27405.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGulabivala K, Ng YL. Factors that affect the outcomes of root canal treatment and retreatment-A reframing of the principles. Int Endod J. 2023;56(Suppl 2):82\u0026ndash;115.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei X, Du Y, Zhou X, Yue L, Yu Q, Hou B, Chen Z, Liang J, Chen W, Qiu L, Huang X, Meng L, Huang D, Wang X, Tian Y, Tang Z, Zhang Q, Miao L, Zhao J, Yang D, Yang J, Ling J. Expert consensus on digital guided therapy for endodontic diseases. Int J Oral Sci. 2023;15(1):54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan R, Akbar S, Khan A, Marwan M, Qaisar ZH, Mehmood A, Shahid F, Munir K, Zheng Z. Dental image enhancement network for early diagnosis of oral dental disease. Sci Rep. 2023;13(1):5312.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu X, Zheng X, Lin F, Yu Q, Hou B, Chen Z, Wei X, Qiu L, Wenxia C, Li J, Chen L, Wang Z, Wu H, Lu Z, Zhao J, Liang Y, Zhao J, Pan Y, Pan S, Wang X, Yang D, Ren Y, Yue L, Zhou X. Expert consensus on endodontic therapy for patients with systemic conditions. Int J Oral Sci. 2024;16(1):45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohammad-Rahimi H, Motamedian SR, Rohban MH, Krois J, Uribe SE, Mahmoudinia E, Rokhshad R, Nadimi M, Schwendicke F. Deep learning for caries detection: A systematic review. J Dent. 2022;122:104115.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSadr S, Mohammad-Rahimi H, Motamedian SR, Zahedrozegar S, Motie P, Vinayahalingam S, Dianat O, Nosrat A. Deep Learning for Detection of Periapical Radiolucent Lesions: A Systematic Review and Meta-analysis of Diagnostic Test Accuracy. J Endod. 2023;49(3):248\u0026ndash;e2613.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Rosa CS, Bergamini ML, Palmieri M, Sarmento DJS, de Carvalho MO, Ricardo ALF, Hasseus B, Jonasson P, Braz-Silva PH, Ferreira Costa AL. Differentiation of periapical granuloma from radicular cyst using cone beam computed tomography images texture analysis. Heliyon. 2020;6(10):e05194.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK\u0026uuml;nzle P, Paris S. Performance of large language artificial intelligence models on solving restorative dentistry and endodontics student assessments. Clin Oral Investig. 2024;28(11):575.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang H, Zheng O, Wang D, Yin J, Wang Z, Ding S, Yin H, Xu C, Yang R, Zheng Q, Shi B. ChatGPT for shaping the future of dentistry: the potential of multi-modal large language model. Int J Oral Sci. 2023;15(1):29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSlayton RL, Urquhart O, Araujo MWB, Fontana M, Guzm\u0026aacute;n-Armstrong S, Nascimento MM, Nov\u0026yacute; BB, Tinanoff N, Weyant RJ, Wolff MS, Young DA, Zero DT, Tampi MP, Pilcher L, Banfield L, Carrasco-Labra A. Evidence-based clinical practice guideline on nonrestorative treatments for carious lesions: A report from the American Dental Association. J Am Dent Assoc. 2018;149(10):837\u0026ndash;e84919.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeters OA, Rossi-Fedele G, George R, Kumar K, Timmerman A, Wright PP. Guidelines for non-surgical root canal treatment. Aust Endod J. 2024;50(2):202\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoubaris M, Cameron A, Manakil J, George R. Artificial intelligence vs. semi-automated segmentation for assessment of dental periapical lesion volume index score: A cone-beam CT study. Comput Biol Med. 2024;175:108527.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTable 1 to 6 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-oral-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ohea","sideBox":"Learn more about [BMC Oral Health](http://bmcoralhealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/ohea/default.aspx","title":"BMC Oral Health","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Large Language Models, Clinical Decision Support, Guideline Adherence, Caries Management","lastPublishedDoi":"10.21203/rs.3.rs-8542862/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8542862/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjective\u003c/h2\u003e \u003cp\u003eTo systematically evaluate the capability of three large language models (LLMs)\u0026mdash;ChatGPT-4o, Grok-3, and DeepSeek\u0026mdash;in interpreting and translating clinical practice guidelines for caries management and in supporting clinical decision-making, thereby exploring their potential role in disseminating dental knowledge and assisting clinical practice.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eBased on the American Dental Association\u0026rsquo;s Evidence-Based Clinical Practice Guideline on Nonrestorative Treatments for Carious Lesions, a zero-shot prompting strategy was used to instruct each model to generate guideline summaries tailored for both healthcare professionals and the general public. Additionally, the models were asked to provide diagnoses and treatment plans for three standardized clinical cases related to caries. Manual evaluations were conducted across five dimensions\u0026mdash;accuracy, clarity, conciseness, logical coherence, and overall quality\u0026mdash;using a 0\u0026ndash;10 scoring system. Text consistency was also assessed using ROUGE-L and BLEU metrics.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn generating guideline summaries for the general public, GPT-4o achieved the highest overall score, excelling particularly in clarity and logical coherence, while DeepSeek performed best in terminology accuracy and fidelity to the source text. For summaries intended for healthcare professionals, all three models performed well, with DeepSeek leading in automated evaluation metrics. In clinical case management, ChatGPT attained the highest composite score, significantly outperforming Grok-3 and DeepSeek, demonstrating superior diagnostic accuracy and clinically relevant treatment recommendations.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eLarge language models show promising potential in translating dental guidelines and assisting clinical decision-making. However, limitations such as insufficient personalization and mechanistic application of guidelines remain. Future efforts should focus on integrating multimodal data, enabling dynamic knowledge updates, and developing human\u0026ndash;AI collaborative care models to achieve a balance between standardized and personalized management of oral diseases.\u003c/p\u003e","manuscriptTitle":"Evaluating Large Language Models for Translating Caries Guidelines into Clinical Decision Support","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-28 03:06:47","doi":"10.21203/rs.3.rs-8542862/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-02-10T12:58:11+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-10T06:37:37+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-27T10:01:00+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"192451263648267504014744325533131805711","date":"2026-01-26T14:05:59+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"168739716933689271138180660055848796698","date":"2026-01-22T13:01:33+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-22T12:08:34+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-01-22T11:20:08+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-12T12:40:46+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-12T12:38:50+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Oral Health","date":"2026-01-07T14:39:23+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-oral-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ohea","sideBox":"Learn more about [BMC Oral Health](http://bmcoralhealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/ohea/default.aspx","title":"BMC Oral Health","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b1081751-54bf-4278-99af-0c52a7e39c6c","owner":[],"postedDate":"January 28th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-05-04T16:05:48+00:00","versionOfRecord":{"articleIdentity":"rs-8542862","link":"https://doi.org/10.1186/s12903-026-08430-3","journal":{"identity":"bmc-oral-health","isVorOnly":false,"title":"BMC Oral Health"},"publishedOn":"2026-04-27 15:58:34","publishedOnDateReadable":"April 27th, 2026"},"versionCreatedAt":"2026-01-28 03:06:47","video":"","vorDoi":"10.1186/s12903-026-08430-3","vorDoiUrl":"https://doi.org/10.1186/s12903-026-08430-3","workflowStages":[]},"version":"v1","identity":"rs-8542862","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8542862","identity":"rs-8542862","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0