Evaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background: This study evaluated the quality of multiple-choice questions (MCQs) generated by ChatGPT-4o compared with faculty written items in periodontology using the Integrated National Board Dental Examination (INBDE) rubric. Methods: Thirty MCQs were assessed in a blinded cross-sectional comparison at Tufts University School of Dental Medicine. 15 questions were generated by ChatGPT-4o based on course objectives and INBDE guidelines, and 15 were randomly selected from the departmental exam bank. Fourteen periodontology faculty members rated each item on six INBDE criteria including clarity, content accuracy, distractor quality, fairness, curricular alignment, and grammar using a five-point Likert scale ranging from poor (1) to excellent (5). Composite scores were analyzed using a generalized linear mixed model. Results: AI-generated items achieved significantly higher composite scores than human-written questions (20.7 ± 4.9 vs 18.3 ± 5.1; p < 0.001). In descriptive comparisons, AI items also received higher ratings across all six domains, particularly in clarity and grammar. Reviewers were unable to reliably identify the source of the items, and 84.1% of AI generated questions were judged suitable for exam use compared with 55.7% of faculty written items. Conclusions: ChatGPT-4o produced high-quality and well-structured MCQs, and reviewers frequently reported difficulty distinguishing their origin in this blinded assessment. While these results highlight the potential value of AI-assisted assessment design, expert supervision remains essential to ensure accuracy, cognitive depth, and alignment with educational standards. AI should be a supportive tool that complements rather than replaces faculty expertise in question development.
Full text 91,376 characters · extracted from preprint-html · click to expand
Evaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study Bushra Ahmad, Livia Valverde, Shruti Jain, Khaled Saleh, Nadeem Karimbux, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8614641/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: This study evaluated the quality of multiple-choice questions (MCQs) generated by ChatGPT-4o compared with faculty written items in periodontology using the Integrated National Board Dental Examination (INBDE) rubric. Methods: Thirty MCQs were assessed in a blinded cross-sectional comparison at Tufts University School of Dental Medicine. 15 questions were generated by ChatGPT-4o based on course objectives and INBDE guidelines, and 15 were randomly selected from the departmental exam bank. Fourteen periodontology faculty members rated each item on six INBDE criteria including clarity, content accuracy, distractor quality, fairness, curricular alignment, and grammar using a five-point Likert scale ranging from poor (1) to excellent (5). Composite scores were analyzed using a generalized linear mixed model. Results: AI-generated items achieved significantly higher composite scores than human-written questions (20.7 ± 4.9 vs 18.3 ± 5.1; p < 0.001). In descriptive comparisons, AI items also received higher ratings across all six domains, particularly in clarity and grammar. Reviewers were unable to reliably identify the source of the items, and 84.1% of AI generated questions were judged suitable for exam use compared with 55.7% of faculty written items. Conclusions: ChatGPT-4o produced high-quality and well-structured MCQs, and reviewers frequently reported difficulty distinguishing their origin in this blinded assessment. While these results highlight the potential value of AI-assisted assessment design, expert supervision remains essential to ensure accuracy, cognitive depth, and alignment with educational standards. AI should be a supportive tool that complements rather than replaces faculty expertise in question development. Dentistry Artificial Intelligence and Machine Learning artificial intelligence periodontology dental education multiple-choice questions educational assessment Figures Figure 1 Figure 2 Figure 3 1. BACKGROUND Multiple-choice questions (MCQs) remain the predominant format for assessing applied knowledge in health professions because they allow objective scoring and efficiently cover wide ranges of course material. [ 1 ] According to established guidelines, a well-constructed MCQ includes a clearly written stem, an unambiguous lead-in, and a set of plausible distractors designed to assess applied knowledge and reduce cueing. [ 2 ] To distinguish higher from lower performers, items must have appropriate difficulty and be challenging enough to differentiate between those who understand the material and those who do not. [ 1 ] Producing such questions requires subject-matter expertise, conceptual integration and meticulous editing to avoid common item-writing flaws such as mutually exclusive distractors, overly long correct answers or absolute terms that hint at the correct option. [ 1 ] Consequently, developing large, high-quality question banks is time-consuming and resource-intensive for faculty. [ 3 ] The rise of artificial intelligence (AI) language models has prompted growing interest in automated exam item generation. Recent advances in natural language processing, particularly through models like ChatGPT-4, have enabled the creation of coherent, grammatically correct, and content-rich test items, supporting their use in efficiently generating educational content and multiple-choice questions. [ 4 ] Preliminary studies across medical education confirm that AI can generate examination items significantly faster than traditional manual methods; however, expert review remains essential to ensure the accuracy, relevance, and overall quality of the generated content. [ 5 , 6 ] For instance, in urology, a majority of questions generated by a customized ChatGPT-4 model for residents demonstrated acceptable or even excellent discriminatory value. [ 2 ] However, performance was inconsistent, with studies in other specialties revealing significant limitations. In emergency medicine, AI-generated questions were significantly easier and tested lower-order cognitive skills, whereas human-authored questions better assessed higher-order clinical analysis. [ 5 ] Similarly, a study in general surgery reported that most AI-generated questions failed to meet acceptable discrimination thresholds, unlike their human-written counterparts. [ 7 ] Factual accuracy remains a concern, as demonstrated by Klang et al. (2023), who found that 15% of AI-generated pathology questions required expert correction. [ 8 ] In the field of radiology, Emekli et al. (2024) evaluated ChatGPT-4o’s performance in generating board-style exam questions and found that, although the model produced coherent text and clinically relevant scenarios, its inability to interpret medical images limited the diagnostic depth and overall quality of the items. [ 9 ] These mixed findings underscore that while AI offers a powerful tool for augmenting assessment, rigorous human oversight is essential for validating accuracy, ensuring clinical relevance, and elevating the cognitive demand of generated questions before their use in medical and dental education. [ 1 , 8 ] Although AI is increasingly used to generate exam questions in other health disciplines, its application to dental specialties, particularly periodontology, remains understudied. Early work evaluating ChatGPT-4 and Bard’s MCQs on narrow topics like dental-caries showed that both models tended to produce low-order questions with limited cognitive depth and occasional flaws in wording or use of absolute terms. [ 10 ] More comprehensive evaluations from a UK dental school, which had ChatGPT-4o, Grok 2 and Gemini generate 140 undergraduate dental questions, found that although the AI-drafted items were coherent, reviewers frequently identified problems such as double negatives, overly long stems and factual inaccuracies. [ 11 ] A randomized trial in periodontology by Ma and colleagues (2025) further highlighted these concerns, finding that while a ChatGPT-4-generated exam led to higher student scores, it was significantly less effective at differentiating between high- and low-performing students, which they attributed to the AI’s inability to create questions with sufficient clinical nuance, a weakness echoed by the students, who, despite scoring higher, rated the AI test as less inspiring and valuable to their learning. [ 4 ] Against this backdrop, in the present study we aimed to evaluate the quality of AI-generated MCQs in periodontology compared with faculty-written items using the Integrated National Board Dental Examination (INBDE) rubric. Although AI can rapidly generate large volumes of dental MCQs, expert review remains essential to ensure that such items meet established standards and maintain clinical relevance. We hypothesized that the AI-generated items would achieve similar or superior quality scores and that reviewers would be unable to reliably identify their origin. Our findings contribute to the growing discourse on AI integration in dental education by offering empirical evidence to guide responsible adoption of generative models in assessment design. 2. METHODS 2.1 Ethics Approval and Consent All procedures involving human participants were conducted in accordance with the ethical standards of the Tufts Health Sciences Institutional Review Board (STUDY00005769), applicable institutional and national regulations, and the principles of the Declaration of Helsinki. All participating faculty provided informed consent and received lunch as the only form of reimbursement for their time. 2.2 Study Design We conducted this single-site, cross-sectional, blinded comparison at Tufts University School of Dental Medicine (TUSDM). We evaluated 30MCQs, with 15 produced by ChatGPT-4o and 15 randomly selected from the 2024-25 predoctoral periodontology examination bank. After the items were merged and randomly ordered, they were uploaded to Qualtrics. In a single evaluation session, faculty reviewers rated every question in the survey platform without knowing whether it came from the human bank or from AI. 2.3 Item Generation For the AI-generated set, we provided ChatGPT-4o with de-identified course objectives and lecture materials from the 2024-25 predoctoral periodontology course along with the INBDE blueprint. We instructed the AI to create 15 content-based MCQs that adhered to INBDE writing guidelines. For the human-generated set, we randomly selected 15 previously vetted MCQs from the faculty-authored examinations for the same 2024-25 course. Both sets of questions addressed the same course materials. At TUSDM, the INBDE rubric serves as the standard framework for constructing and reviewing MCQs in departmental examinations. However, because faculty authors independently develop their items, adherence to every INBDE criterion cannot be guaranteed. We exported all 30 questions to Qualtrics for blind review. 2.4 Expert Reviewers Of the 22 full- and part-time faculty members in the Department of Periodontology, 14completed the evaluation, a response rate of 63.6%. Before rating began, reviewers attended a 20-minute calibration session that covered the INBDE item-quality rubric, demonstrated common flaws, and provided an opportunity to ask clarifying questions. 2.5 Evaluation Instrument Each question was assessed on six INBDE-based criteria, including clarity, content accuracy, distractor quality, fairness, curricular alignment, and grammar, using a five-point Likert scale ranging from Poor (1) to Excellent (5). After scoring the six domains, reviewers answered two yes-or-no questions: whether they believed the item had been developed by AI or written by a human, and whether the item was suitable for inclusion on a summative examination. 2.6 Data Handling and Composite Score To manage sparsely populated categories, Likert responses were collapsed so that scores of 1 and 2 were combined, score 3 was left unchanged, and scores of 4 and 5 were combined. A composite quality score ranging from 0 to 24 was then calculated for each item by summing the recoded values for the six criteria. 2.7 Statistical Analysis Descriptive statistics were generated for each criterion and for the composite score, both overall and stratified by item source. Differences in composite score between AI and human items by using a generalized linear mixed model to account for the clustering of different criteria for each question. A two-sided p-value of 0.05 or lower was considered statistically significant. All analyses were performed using SAS Version 9.4 (SAS Institute Inc., Cary, NC, USA). 3. RESULTS 3.1 Participant Characteristics Fourteen periodontology faculty members at Tufts University completed the blinded evaluation of all 30 MCQs, yielding a participation rate of 63.6%. Each participant rated 15 human‑written questions randomly sampled from the TUSDM predoctoral periodontology departmental bank and fifteen questions generated by ChatGPT‑4o; question order was randomized. After collapsing the five-point Likert ratings into three categories and summing the six INBDE criteria, the composite quality scores ranged from 0 to 24. A total of 420 observations were available for analysis, with 196 ratings of human items and 224 ratings of AI items. 3.2 Domain‑Level Ratings 3.2.1 Question Clarity On the clarity criterion, human‑written items exhibited a relatively even distribution across rating categories: 31.4% of ratings fell in the poorest category (1), 34.5% were moderate (3) and 34.0% were rated good/excellent (4). By contrast, AI‑generated items received high clarity scores in the majority of evaluations: only 8.7% were rated poor, 28.8% were moderate and 62.6% were rated good/excellent. In this sample, AI items were more frequently judged clearly written than human items. 3.2.2 Content Accuracy For the content accuracy criterion, 44.0% of ratings for human-written items were in the good/excellent category and 22.3% were rated poor. In comparison, AI-generated questions had a higher proportion of good/excellent ratings in this sample (64.8% vs. 44.0%) and a smaller proportion of poor ratings (6.4% vs. 22.3%). The proportion of moderate ratings was similar between groups (33.7% human vs. 28.8% AI). 3.2.3 Quality of Distractors ChatGPT-4o often produced more plausible distractors than did the faculty. Human‑written items again showed a broad distribution: 29.4% were rated poor, 30.4% moderate, and 40.2% good/excellent. AI items were judged substantially higher, with 63.0% of ratings in the good/excellent category and only 15.5% rated poor. These findings indicate that ChatGPT‑4o often produced more plausible distractors than did the faculty. 3.2.4 Fairness Both groups achieved high fairness scores. Human items were rated excellent in 66.3% of observations and poor in 7.3%, whereas AI items were rated excellent in 71.7% and poor in 2.3%. Fairness ratings were high for both groups, with a slightly higher proportion of AI items rated excellent in this sample. 3.2.5 Alignment with the Periodontology Curriculum Curricular alignment scores revealed that 58.8% of human questions and 72.6% of AI questions were well- aligned with the periodontology curriculum. Poor alignment was noted in 8.3% of human items compared to only 3.6% of AI items. 3.2.6 Grammar and Language Human‑written items were rated grammatically sound in 46.1% of observations, moderate in 35.2% and poor in 18.7%. ChatGPT‑4o‑generated questions demonstrated superior linguistic quality, with 68.8% rated good/excellent and only 6.4% rated poor. 3.3 Ability to Identify Item Origin After rating each question, reviewers guessed whether it was generated by AI or by a human writer. Their ability to identify the source was poor (Fig. 1 ). For human-written items, only 37.8% of guesses correctly indicated human origin, while 33.2% incorrectly labelled them as AI and 29.0% selected “cannot tell”. For AI-generated questions, just 27.5% of guesses correctly indicated AI origin, 41.4% misclassified them as human, and 31.1% were unsure. These results suggest that faculty reviewers could not reliably distinguish AI-generated from human-written items. 3.4 Suitability for Exam Use Reviewers also indicated whether each question was suitable for inclusion in a periodontology exam (Fig. 2 ). Among the human-written items, 55.7% were considered suitable for inclusion in a periodontology exam, 23.2% were not suitable, and 21.1% were considered acceptable only with modifications. By comparison, 84.1% of AI-generated items were judged suitable, 6.8% were not suitable, and 9.1% might be usable with modification. Thus, faculty were far more likely to recommend AI questions for direct examination use. 3.5 Composite Quality Scores The composite score, derived by summing participant responses to the six recoded criteria (range 0–24), provided an overall measure of item quality (Fig. 3 ). Human-written questions had a mean ± SD composite score of 18.3 ± 5.1 with a median of 18.0 and an interquartile range (IQR) of 6.0. AI-generated items achieved a higher mean ± SD score of 20.7 ± 4.9 with a median of 23.0 and the same IQR of 6.0. A generalized linear mixed model treating the question as a random effect and the source as a fixed effect indicated that AI generation increased the composite score by 2.38 points (standard error 0.54; t = − 4.43; p < 0.0001). 3.6 Criterion‑Level “Excellent” Ratings Because the INBDE rubric was collapsed into three categories (poor = 1, moderate = 3, excellent = 4), we compared the proportion of responses rated “excellent” for each quality criterion (Table 1 ). Across all six criteria, ChatGPT‑generated questions were more likely to receive top ratings, with the largest absolute advantage in clarity (≈ 29 percentage points) and the smallest in fairness (≈ 5 points). These findings indicate that AI items were frequently rated in the highest category across multiple criteria in this sample. Table 1 Proportion of “Excellent” Ratings Across INBDE Quality Criteria for Human- and AI-Generated Questions. Criterion Human “excellent” (%) ChatGPT “excellent” (%) Difference (percentage points) * Question clarity 34.0 62.6 + 28.6 Content accuracy 44.0 64.8 + 20.8 Quality of distractors 40.2 63.0 + 22.8 Fairness 66.3 71.7 + 5.4 Alignment with curriculum 58.8 72.6 + 13.8 Grammar & language 46.1 68.8 + 22.7 *Percentage-point difference relative to the human-generated group, which was considered the reference. 4. DISCUSSION In this study we evaluated the quality of AI-generated MCQs in Periodontology, comparing items created by ChatGPT-4o with those written by faculty, using the INBDE rubric as a standardized framework. Our findings revealed that AI-generated questions had a significantly higher composite quality score and, descriptively, received higher ratings across the six evaluated domains, and they were also more frequently deemed suitable for inclusion in summative examinations. Reviewers were also largely unable to determine whether the questions were produced by AI or by human authors. Collectively, these results suggest that ChatGPT-4o can generate high-quality, well-structured assessment items that are indistinguishable from, and sometimes superior to, traditional faculty-written questions. One possible explanation is that large language models can adhere strictly to explicit instructions, formatting rules, and item-writing guidelines, producing questions that are more consistently structured than those written by humans, who may vary in approach or inadvertently deviate from recommended standards. The significantly higher composite score for AI-generated items aligns with findings from several recent evaluations of LLMs in health-professions assessment. Artsi et al. (2024), for example, compared ChatGPT-4, ChatGPT-3.5, and Med-PaLM 2 in generating USMLE-style items and found that all three models produced coherent, blueprint-aligned questions, with ChatGPT-4 yielding the highest editorial quality. [ 1 ] In emergency medicine, Law et al. (2023) reported that ChatGPT-4o generated MCQs were grammatically strong but significantly easier and less discriminative than expert-written items. [ 5 ] Ahmed et al. (2025) showed that ChatGPT and Bard could produce linguistically acceptable dental-caries MCQs, although many required refinements for depth, and Dave et al. (2025) found that AI-platforms as ChatGPT-4o, Grok 2, and Gemini, generated coherent undergraduate dentistry items but sometimes included double negatives, overly long stems, or factual errors. [ 12 , 11 ] In Periodontology, Ma et al. (2025) demonstrated that ChatGPT-4-generated exams produced higher student scores but poorer discrimination between high- and low-performers, suggesting insufficient cognitive challenge. [ 4 ] In contrast to these mixed findings, our results show that ChatGPT-4o outperformed human-written items across all six INBDE domains and was more often judged suitable for exam use. However, consistent with the broader literature, we also did not evaluate psychometric performance, so this observed higher editorial quality cannot yet be interpreted as evidence of greater cognitive rigor. In our study, AI-generated items had higher proportions of top ratings for clarity, content accuracy, and curricular alignment in descriptive comparisons. Similar findings have been reported elsewhere. For example, Artsi et al. (2024) showed that ChatGPT-4 produced USMLE-style items with fewer structural flaws than human authors, and Cheung et al. (2023) found that ChatGPT-based urology questions were often more grammatically polished. [ 1 , 6 ] Dental studies, including Dave et al. (2025), likewise noted that models such as ChatGPT-4o could generate items that read more clearly than faculty drafts. [ 11 ] However, these same investigations also highlighted that highly polished items are not always conceptually accurate or fully on-topic. Thus, while structured prompts and context scaffolding likely contributed to the strong performance we observed, expert review remains essential to ensure that fidelity to the intended curriculum. The superior ratings likely reflect the model’s ability to adhere to established principles for MCQs, which demand a clear stem, a focused lead-in, and plausible distractors of similar length and grammatical structure. [ 13 ] Common pitfalls include grammatical or logical cues, use of extreme terms (“always”, “never”), and distractors that are obviously wrong. [ 14 ] In our study, the AI consistently avoided these flaws and provided balanced distractors, demonstrating a level of rule-following and consistency that is difficult for human authors to maintain over large item sets. This ability to consistently follow item-writing guidelines may contribute to the higher composite quality ratings observed for AI-generated items. However, high-quality questions also require alignment with learning outcomes and validity evidence, such as item difficulty and discrimination indices. [ 13 ] Because we did not administer the questions to students, psychometric qualities remain unknown. Post-hoc analysis of item statistics and internal structure will be needed to confirm whether the higher expert ratings translate into reliable discrimination among learners. [ 14 ] An additional finding of this study was that reviewers could not reliably determine whether an item had been written by AI or by a human. This suggests that AI-generated questions now reach a level of fluency and structure that closely resembles expert authorship [ 12 ]. From an academic-integrity standpoint, this indistinguishability highlights the need for clear safeguards, because perceived authorship can no longer serve as an indicator of quality. Rigorous, rubric-based review should therefore remain standard practice for all items, regardless of their source. The fact that AI-generated questions were more often judged immediately suitable for exam use also suggests that this technology may serve as a useful supplement to faculty-driven assessment development. The implications for dental education are significant. Integrating AI into assessment workflows could help address one of the most persistent challenges in curriculum design: the shortage of time and faculty resources for generating new, high-quality test items [ 15 ]. Timing advantages have been clearly demonstrated; for example, Cheung et al. (2023) reported that their ChatGPT-based system generated urology MCQs in a fraction of the time required for faculty authors, while maintaining acceptable editorial quality, often within seconds. [ 6 ] With LLMs capable of producing questions that meet or exceed linguistic standards, educators may shift their focus toward validation, cognitive-level enhancement, and psychometric calibration rather than first-draft item creation [ 11 ]. This could enable the development of larger, more diverse, and continuously updated question banks. Our study’s design offered several notable strengths. It featured a fully blinded, head-to-head comparison that applied all six INBDE quality criteria. Fourteen evaluators participated, drawn from both full and part-time faculty, all of whom are American Board certified by the American Academy of Periodontology, which enhanced ecological validity. The study also reported both criterion level and composite outcomes to provide a more nuanced picture of item performance. Despite these encouraging results, the study has several limitations. The work was conducted within a single institution and limited to the discipline of Periodontology, which may restrict the generalizability of the findings to other educational contexts. Additionally, the outcomes were based exclusively on expert ratings; we did not collect psychometric data such as item difficulty or student performance. Finally, the results are specific to a particular large language model, ChatGPT-4o, and the prompting approach used, both of which may influence item quality. Several authors have cautioned against relying too heavily on AI for summative assessments without strong faculty involvement. Norcini et al. (2018) note that high-stakes exams still require careful blueprinting and psychometric review after administration, which are processes that remain fundamentally human-driven. [ 16 ] Others have highlighted the risk of construct underrepresentation, where AI-generated items may focus mainly on lower-level recall unless they are intentionally designed to assess higher-order thinking. [ 17 , 18 ] These concerns are consistent with earlier findings showing that AI-generated questions, although polished, may have lower discriminative ability if used without iterative refinement. From an implementation standpoint, best-practice frameworks increasingly advocate for a hybrid model in which AI serves as an initial drafting tool, followed by structured expert review and psychometric validation. [ 18 ] Within such a model, AI-generated items may enhance efficiency without compromising assessment integrity. Our findings support this approach, suggesting that while AI can substantially improve the efficiency and consistency of item development, its optimal use lies in augmentation rather than replacement of faculty expertise. Establishing clear governance structures, reviewer training, and documentation of AI involvement will be essential as these tools become more deeply integrated into dental education. Future research should include larger, multicenter samples across multiple specialties, embed AI-generated items in live examinations to capture psychometric properties, and compare different LLMs or retrieval-augmented architectures. Studies should also assess the time and cost savings of AI-assisted item generation, evaluate learner acceptance, and explore whether fine-tuned models further improve quality and cognitive level. 5. CONCLUSIONS This study demonstrates that ChatGPT-4o can generate periodontology MCQs that achieved a significantly higher composite quality score in this blinded evaluation. Reviewers were largely unable to distinguish AI-generated from human-authored questions, and most AI items were deemed suitable for direct examination use. These findings suggest that LLMs can serve as valuable adjuncts in creating dental assessment materials, helping educators expand and refresh question banks with remarkable efficiency. However, while the editorial and structural quality of AI-generated questions was strong, this does not necessarily translate into higher cognitive rigor or validity in discriminating learner performance. Expert oversight, therefore, remains essential to verify factual accuracy, capture clinical nuances, and maintain appropriate cognitive demand. Rather than replacing human judgment, AI should be viewed as a supportive tool that can streamline the drafting process while faculty remain central to content validation and educational integrity. Future work should integrate AI-generated questions into live assessments to examine psychometric performance and explore strategies to optimize collaboration between educators and generative models in dental education. Declarations Ethics Approval and Consent to Participate The study protocol received approval from the Tufts Health Sciences Institutional Review Board (STUDY00005769) and was classified as minimal‑risk research. Consent for Publication All authors in this research provide consent for publication. Availability of Data and Materials The datasets supporting the conclusions of this article are available from the corresponding author upon reasonable request. Competing Interests All authors in this research declare no competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Authors’ Contributions BA designed the study, acquired the data, and wrote the manuscript. LV acquired the data and revised the manuscript. SJ interpreted the data and performed the data analysis. KS wrote the manuscript. NK critically revised the manuscript. NJ critically revised the manuscript. All authors read and approved the final manuscript. Acknowledgments Not applicable. References Artsi M, Karras BT, Palanica A. Evaluating the quality of AI-generated USMLE-style exam questions across ChatGPT-4, GPT-3.5, and Med-PaLM 2. BMJ Open. 2024;14:e082559. doi:10.1136/bmjopen-2023-082559. Kim JK, Kim SJ, Park JY, Ahn HS. Evaluating ChatGPT-4 in generating and assessing multiple-choice questions in urology. Int J Med Educ. 2025;16:23–29. doi:10.5116/ijme.6623.faf2 Abouzeid E, Wassef R, Jawwad A, Harris P. Chatbots’ role in generating single best answer questions for undergraduate medical student assessment: comparative analysis. JMIR Med Educ. 2025;11:e69521. doi:10.2196/69521 Ma X, Pan W, Yu X. Evaluating AI-generated examination papers in periodontology: a comparative study with human-designed counterparts. BMC Med Educ. 2025;25:1099. doi:10.1186/s12909-025-07706-6 Law AKK, Law KIW, Law MF, Lee KKC. Evaluation of multiple-choice questions generated by ChatGPT in emergency medicine: comparative cross-sectional study. JMIR Med Educ. 2023;9:e49174. doi:10.2196/49174 Cheung BHH, Sim JJL, Lam ST, Wong RHL, Lai VWY. Evaluating the effectiveness and validity of artificial intelligence–generated urology multiple-choice questions. Adv Med Educ Pract. 2023;14:1227–1235. doi:10.2147/AMEP.S423138 Kıyak YS, Coşkun AK, Kaymak Ş, Coşkun Ö, Budakoğlu İİ. Can ChatGPT generate surgical multiple-choice questions comparable to those written by a surgeon? Proc (Bayl Univ Med Cent). 2024;38(1):48–52. doi:10.1080/08998280.2024.2418752 Klang E, Perelman V, Avnon T, Klang S, Halpern P. Utility of ChatGPT in multiple choice question writing for medical students. BMC Med Educ. 2023;23:908. doi:10.1186/s12909-023-04669-1 Emekli E, Yıldırım S, Aydın E. Can ChatGPT generate appropriate radiology multiple-choice questions for medical students? A pilot study. Clin Imaging. 2024;98:44–49. doi:10.1016/j.clinimag.2023.11.005 Ahmed WM, Azhari AA, Alfaraj A, Alhamadani A, Zhang M, Lu CT. The quality of AI-generated dental caries multiple-choice questions: a comparative analysis of ChatGPT and Google Bard language models. Heliyon. 2024;10:e28198. doi:10.1016/j.heliyon.2024.e28198 Dave M, Tattar R, Alafaleg R, Barry S, Ariyaratnam S, Roudsari RV, et al. Performance of large language models (ChatGPT-4o, Grok-2, and Gemini) in UK dentistry and dental hygiene and therapy assessments. Br Dent J. 2025. doi:10.1038/s41415-025-8383-2 Ahmed A, Kerr E, O’Malley A. Quality assurance and validity of AI-generated single best answer questions. BMC Med Educ. 2025;25:300. doi:10.1186/s12909-025-06881-w Gottlieb M, Bailitz J, Fix M, Shappell E. Educator’s blueprint: a how-to guide for developing high-quality multiple-choice questions. AEM Educ Train. 2023;7(1):e10836. doi:10.1002/aet2.10836 Naidoo M. The pearls and pitfalls of setting high-quality multiple choice questions for clinical medicine. S Afr Fam Pract. 2023;65(1):e1–e4. doi:10.4102/safp.v65i1.5726 Dunlap CA, Lai G. Enhancing dental education through AI: revolutionizing exam creation and grading [Internet]. American Association of Endodontists; 2024 Dec 20 [cited 2026 Jan 2]. Available from: https://www.aae.org/specialty/enhancing-dental-education-through-ai-revolutionizing-exam-creation-and-grading/ Norcini J, Anderson MB, Bollela V, et al. 2018 consensus framework for good assessment. Med Teach . 2018;40(11):1102–1109. Bloom BS, Engelhart MD, Furst EJ, Hill WH, Krathwohl DR. Taxonomy of educational objectives: the classification of educational goals . New York (NY): Longmans; 1956. Schuwirth LWT, van der Vleuten CPM. Programmatic assessment: from assessment of learning to assessment for learning. Med Teach . 2011;33(6):478–485. Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8614641","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":575328062,"identity":"c7947f55-a1e0-47c1-8084-2e6e892d7247","order_by":0,"name":"Bushra Ahmad","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/klEQVRIiWNgGAWjYBACewYGNhDNYwAkJD4wMPCDhXnwaDFsgGiRA2mRnJHAINlDSIvBAYgWY5AWaR5itBjOSH/2uKKGIXG72OGHt21/HJaw5z/A+OBtGx6/SOSYG545xpC4c3aasXVOwmEJHokEZsO5eLQYzshhk2xgY0jccDvBTBqopY5HgoFNmhePFoMb6c8kG/4x1G+4nf5N2gJkC/8B9t/4tSSYSTa2gWzJMZNmAGlhSGBjxqfFsOcNUEufBEhLsWVPWroEz43EZsk55/B4nx3ksG82QC3pG2/8sLGWYO8/fPDDmzLcWqBAApnD2EBQ/SgYBaNgFIwC/AAAAl9Pza9i1tcAAAAASUVORK5CYII=","orcid":"https://orcid.org/0009-0003-0376-0996","institution":"Tufts University School of Dental Medicine","correspondingAuthor":true,"prefix":"","firstName":"Bushra","middleName":"","lastName":"Ahmad","suffix":""},{"id":575328063,"identity":"9918ba1c-1850-440d-9dc1-f038cb4fad80","order_by":1,"name":"Livia Valverde","email":"","orcid":"","institution":"Tufts University School of Dental Medicine","correspondingAuthor":false,"prefix":"","firstName":"Livia","middleName":"","lastName":"Valverde","suffix":""},{"id":575328064,"identity":"2efbae26-6e0f-48bf-b121-7d76e791555b","order_by":2,"name":"Shruti Jain","email":"","orcid":"","institution":"Tufts University School of Dental Medicine","correspondingAuthor":false,"prefix":"","firstName":"Shruti","middleName":"","lastName":"Jain","suffix":""},{"id":575328065,"identity":"de28e236-6af4-4709-aa53-1cac57e35411","order_by":3,"name":"Khaled Saleh","email":"","orcid":"https://orcid.org/0009-0000-2716-5786","institution":"University of Detroit Mercy School of Dentistry","correspondingAuthor":false,"prefix":"","firstName":"Khaled","middleName":"","lastName":"Saleh","suffix":""},{"id":575328066,"identity":"48b3fd50-fdc9-4672-acb3-2cd4bb59d5d6","order_by":4,"name":"Nadeem Karimbux","email":"","orcid":"","institution":"Tufts University School of Dental Medicine","correspondingAuthor":false,"prefix":"","firstName":"Nadeem","middleName":"","lastName":"Karimbux","suffix":""},{"id":575328067,"identity":"acd7715b-f3ed-4207-9d46-fe660e7b6eae","order_by":5,"name":"Y. Natalie Jeong","email":"","orcid":"","institution":"Tufts University School of Dental Medicine","correspondingAuthor":false,"prefix":"","firstName":"Y.","middleName":"Natalie","lastName":"Jeong","suffix":""}],"badges":[],"createdAt":"2026-01-16 02:47:17","currentVersionCode":1,"declarations":{"humanSubjects":true,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":true,"humanSubjectConsent":true,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-8614641/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8614641/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100545372,"identity":"c0190e2a-ccee-417b-93ea-6c57b78ef3cb","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":118796,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/56bd0e6539f8295477426cca.docx"},{"id":100545371,"identity":"9fd1dfc5-ba13-4a7f-bd79-adc59eaa47e3","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs8614641.json","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/0a6daaeac297300faa676739.json"},{"id":100545376,"identity":"40dc1aa9-729f-441f-bafc-33c1bef0843e","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":70205,"visible":true,"origin":"","legend":"","description":"","filename":"rs86146410enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/39756d76812910e5907ceb9b.xml"},{"id":100545383,"identity":"a3b451a2-f7e5-4f00-a345-30d54363e5ff","added_by":"auto","created_at":"2026-01-19 06:28:21","extension":"jpeg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":23929,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/81819fdd9f6830234685aed3.jpeg"},{"id":100545373,"identity":"46b463d6-b91b-40eb-89b2-8379a9623491","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26562,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/745317711f80bc5d8722ea20.jpeg"},{"id":100545377,"identity":"1118ace9-3d63-4557-92b3-d0baf7c211cb","added_by":"auto","created_at":"2026-01-19 06:28:20","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18479,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/5f3be4df9b6943b15990363d.jpeg"},{"id":100548848,"identity":"560dfd8e-5ab5-4074-8178-ebf6f95ba735","added_by":"auto","created_at":"2026-01-19 08:21:13","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7726,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/eec329537c291e3592b9d0be.png"},{"id":100549469,"identity":"20844c66-e6df-4cef-9678-af5ae96795dd","added_by":"auto","created_at":"2026-01-19 08:23:24","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":8786,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/f0daa0378490f31782907da9.png"},{"id":100545374,"identity":"1ab81d2d-dc05-47c3-bc12-8a0aaccf3705","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6777,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/05bb94221dbbf4c7a321607c.png"},{"id":100548916,"identity":"e6e58447-cd42-4970-a750-07a0c95e073a","added_by":"auto","created_at":"2026-01-19 08:21:32","extension":"xml","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":67917,"visible":true,"origin":"","legend":"","description":"","filename":"rs86146410structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/bb391389effd126b6629eb38.xml"},{"id":100549233,"identity":"8120794c-6306-4538-bc11-14fa9fc4a4ab","added_by":"auto","created_at":"2026-01-19 08:22:51","extension":"html","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":76185,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/3a6fb7aaae804c1fca238a9e.html"},{"id":100545370,"identity":"4d2f463a-8ea6-4cbb-abc0-0fd19aada7b6","added_by":"auto","created_at":"2026-01-19 06:28:19","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":23929,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eReviewer Ability to Identify Item Origin (AI vs. Human-Written Questions).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/ce9d697c340f466ec8fbe38d.jpg"},{"id":100548728,"identity":"5c2d92a5-619e-4a26-b6b5-fdbcd1528efa","added_by":"auto","created_at":"2026-01-19 08:20:43","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":26562,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eReviewer Judgements of Examination Suitability for AI-Generated vs. Human-Written Items.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/d98b5d6bca5ea88122eb4e76.jpg"},{"id":100549004,"identity":"f0657463-484a-40f5-8259-cc38d483a3b4","added_by":"auto","created_at":"2026-01-19 08:21:59","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":18479,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution of Composite Quality Scores for AI-Generated and Human-Written Items.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/346b24ad3e7d70a4409c224c.jpg"},{"id":100595107,"identity":"86bcd7f5-5848-414a-ab39-820121cbc2c3","added_by":"auto","created_at":"2026-01-19 13:47:26","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":913013,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8614641/v1/9ce0578e-29b0-48eb-86ec-808427eee048.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eEvaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. BACKGROUND","content":"\u003cp\u003eMultiple-choice questions (MCQs) remain the predominant format for assessing applied knowledge in health professions because they allow objective scoring and efficiently cover wide ranges of course material. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] According to established guidelines, a well-constructed MCQ includes a clearly written stem, an unambiguous lead-in, and a set of plausible distractors designed to assess applied knowledge and reduce cueing. [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] To distinguish higher from lower performers, items must have appropriate difficulty and be challenging enough to differentiate between those who understand the material and those who do not. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] Producing such questions requires subject-matter expertise, conceptual integration and meticulous editing to avoid common item-writing flaws such as mutually exclusive distractors, overly long correct answers or absolute terms that hint at the correct option. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] Consequently, developing large, high-quality question banks is time-consuming and resource-intensive for faculty. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThe rise of artificial intelligence (AI) language models has prompted growing interest in automated exam item generation. Recent advances in natural language processing, particularly through models like ChatGPT-4, have enabled the creation of coherent, grammatically correct, and content-rich test items, supporting their use in efficiently generating educational content and multiple-choice questions. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] Preliminary studies across medical education confirm that AI can generate examination items significantly faster than traditional manual methods; however, expert review remains essential to ensure the accuracy, relevance, and overall quality of the generated content. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] For instance, in urology, a majority of questions generated by a customized ChatGPT-4 model for residents demonstrated acceptable or even excellent discriminatory value. [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] However, performance was inconsistent, with studies in other specialties revealing significant limitations.\u003c/p\u003e \u003cp\u003eIn emergency medicine, AI-generated questions were significantly easier and tested lower-order cognitive skills, whereas human-authored questions better assessed higher-order clinical analysis. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] Similarly, a study in general surgery reported that most AI-generated questions failed to meet acceptable discrimination thresholds, unlike their human-written counterparts. [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] Factual accuracy remains a concern, as demonstrated by Klang et al. (2023), who found that 15% of AI-generated pathology questions required expert correction. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] In the field of radiology, Emekli et al. (2024) evaluated ChatGPT-4o\u0026rsquo;s performance in generating board-style exam questions and found that, although the model produced coherent text and clinically relevant scenarios, its inability to interpret medical images limited the diagnostic depth and overall quality of the items. [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThese mixed findings underscore that while AI offers a powerful tool for augmenting assessment, rigorous human oversight is essential for validating accuracy, ensuring clinical relevance, and elevating the cognitive demand of generated questions before their use in medical and dental education. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eAlthough AI is increasingly used to generate exam questions in other health disciplines, its application to dental specialties, particularly periodontology, remains understudied. Early work evaluating ChatGPT-4 and Bard\u0026rsquo;s MCQs on narrow topics like dental-caries showed that both models tended to produce low-order questions with limited cognitive depth and occasional flaws in wording or use of absolute terms. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] More comprehensive evaluations from a UK dental school, which had ChatGPT-4o, Grok 2 and Gemini generate 140 undergraduate dental questions, found that although the AI-drafted items were coherent, reviewers frequently identified problems such as double negatives, overly long stems and factual inaccuracies. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] A randomized trial in periodontology by Ma and colleagues (2025) further highlighted these concerns, finding that while a ChatGPT-4-generated exam led to higher student scores, it was significantly less effective at differentiating between high- and low-performing students, which they attributed to the AI\u0026rsquo;s inability to create questions with sufficient clinical nuance, a weakness echoed by the students, who, despite scoring higher, rated the AI test as less inspiring and valuable to their learning. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eAgainst this backdrop, in the present study we aimed to evaluate the quality of AI-generated MCQs in periodontology compared with faculty-written items using the Integrated National Board Dental Examination (INBDE) rubric. Although AI can rapidly generate large volumes of dental MCQs, expert review remains essential to ensure that such items meet established standards and maintain clinical relevance. We hypothesized that the AI-generated items would achieve similar or superior quality scores and that reviewers would be unable to reliably identify their origin. Our findings contribute to the growing discourse on AI integration in dental education by offering empirical evidence to guide responsible adoption of generative models in assessment design.\u003c/p\u003e"},{"header":"2. METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Ethics Approval and Consent\u003c/h2\u003e \u003cp\u003e All procedures involving human participants were conducted in accordance with the ethical standards of the Tufts Health Sciences Institutional Review Board (STUDY00005769), applicable institutional and national regulations, and the principles of the Declaration of Helsinki. All participating faculty provided informed consent and received lunch as the only form of reimbursement for their time.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Study Design\u003c/h2\u003e \u003cp\u003eWe conducted this single-site, cross-sectional, blinded comparison at Tufts University School of Dental Medicine (TUSDM). We evaluated 30MCQs, with 15 produced by ChatGPT-4o and 15 randomly selected from the 2024-25 predoctoral periodontology examination bank. After the items were merged and randomly ordered, they were uploaded to Qualtrics. In a single evaluation session, faculty reviewers rated every question in the survey platform without knowing whether it came from the human bank or from AI.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Item Generation\u003c/h2\u003e \u003cp\u003eFor the AI-generated set, we provided ChatGPT-4o with de-identified course objectives and lecture materials from the 2024-25 predoctoral periodontology course along with the INBDE blueprint. We instructed the AI to create 15 content-based MCQs that adhered to INBDE writing guidelines. For the human-generated set, we randomly selected 15 previously vetted MCQs from the faculty-authored examinations for the same 2024-25 course. Both sets of questions addressed the same course materials. At TUSDM, the INBDE rubric serves as the standard framework for constructing and reviewing MCQs in departmental examinations. However, because faculty authors independently develop their items, adherence to every INBDE criterion cannot be guaranteed. We exported all 30 questions to Qualtrics for blind review.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Expert Reviewers\u003c/h2\u003e \u003cp\u003eOf the 22 full- and part-time faculty members in the Department of Periodontology, 14completed the evaluation, a response rate of 63.6%. Before rating began, reviewers attended a 20-minute calibration session that covered the INBDE item-quality rubric, demonstrated common flaws, and provided an opportunity to ask clarifying questions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Evaluation Instrument\u003c/h2\u003e \u003cp\u003eEach question was assessed on six INBDE-based criteria, including clarity, content accuracy, distractor quality, fairness, curricular alignment, and grammar, using a five-point Likert scale ranging from Poor (1) to Excellent (5). After scoring the six domains, reviewers answered two yes-or-no questions: whether they believed the item had been developed by AI or written by a human, and whether the item was suitable for inclusion on a summative examination.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Data Handling and Composite Score\u003c/h2\u003e \u003cp\u003eTo manage sparsely populated categories, Likert responses were collapsed so that scores of 1 and 2 were combined, score 3 was left unchanged, and scores of 4 and 5 were combined. A composite quality score ranging from 0 to 24 was then calculated for each item by summing the recoded values for the six criteria.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7 Statistical Analysis\u003c/h2\u003e \u003cp\u003eDescriptive statistics were generated for each criterion and for the composite score, both overall and stratified by item source. Differences in composite score between AI and human items by using a generalized linear mixed model to account for the clustering of different criteria for each question. A two-sided p-value of 0.05 or lower was considered statistically significant. All analyses were performed using SAS Version 9.4 (SAS Institute Inc., Cary, NC, USA).\u003c/p\u003e \u003c/div\u003e"},{"header":"3. RESULTS","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Participant Characteristics\u003c/h2\u003e \u003cp\u003eFourteen periodontology faculty members at Tufts University completed the blinded evaluation of all 30 MCQs, yielding a participation rate of 63.6%. Each participant rated 15 human‑written questions randomly sampled from the TUSDM predoctoral periodontology departmental bank and fifteen questions generated by ChatGPT‑4o; question order was randomized. After collapsing the five-point Likert ratings into three categories and summing the six INBDE criteria, the composite quality scores ranged from 0 to 24. A total of 420 observations were available for analysis, with 196 ratings of human items and 224 ratings of AI items.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Domain‑Level Ratings\u003c/h2\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Question Clarity\u003c/h2\u003e \u003cp\u003eOn the clarity criterion, human‑written items exhibited a relatively even distribution across rating categories: 31.4% of ratings fell in the poorest category (1), 34.5% were moderate (3) and 34.0% were rated good/excellent (4). By contrast, AI‑generated items received high clarity scores in the majority of evaluations: only 8.7% were rated poor, 28.8% were moderate and 62.6% were rated good/excellent. In this sample, AI items were more frequently judged clearly written than human items.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2 Content Accuracy\u003c/h2\u003e \u003cp\u003eFor the content accuracy criterion, 44.0% of ratings for human-written items were in the good/excellent category and 22.3% were rated poor. In comparison, AI-generated questions had a higher proportion of good/excellent ratings in this sample (64.8% vs. 44.0%) and a smaller proportion of poor ratings (6.4% vs. 22.3%). The proportion of moderate ratings was similar between groups (33.7% human vs. 28.8% AI).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3 Quality of Distractors\u003c/h2\u003e \u003cp\u003eChatGPT-4o often produced more plausible distractors than did the faculty. Human‑written items again showed a broad distribution: 29.4% were rated poor, 30.4% moderate, and 40.2% good/excellent. AI items were judged substantially higher, with 63.0% of ratings in the good/excellent category and only 15.5% rated poor. These findings indicate that ChatGPT‑4o often produced more plausible distractors than did the faculty.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e3.2.4 Fairness\u003c/h2\u003e \u003cp\u003eBoth groups achieved high fairness scores. Human items were rated excellent in 66.3% of observations and poor in 7.3%, whereas AI items were rated excellent in 71.7% and poor in 2.3%. Fairness ratings were high for both groups, with a slightly higher proportion of AI items rated excellent in this sample.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003e3.2.5 Alignment with the Periodontology Curriculum\u003c/h2\u003e \u003cp\u003eCurricular alignment scores revealed that 58.8% of human questions and 72.6% of AI questions were well- aligned with the periodontology curriculum. Poor alignment was noted in 8.3% of human items compared to only 3.6% of AI items.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e3.2.6 Grammar and Language\u003c/h2\u003e \u003cp\u003eHuman‑written items were rated grammatically sound in 46.1% of observations, moderate in 35.2% and poor in 18.7%. ChatGPT‑4o‑generated questions demonstrated superior linguistic quality, with 68.8% rated good/excellent and only 6.4% rated poor.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Ability to Identify Item Origin\u003c/h2\u003e \u003cp\u003eAfter rating each question, reviewers guessed whether it was generated by AI or by a human writer. Their ability to identify the source was poor (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For human-written items, only 37.8% of guesses correctly indicated human origin, while 33.2% incorrectly labelled them as AI and 29.0% selected \u0026ldquo;cannot tell\u0026rdquo;. For AI-generated questions, just 27.5% of guesses correctly indicated AI origin, 41.4% misclassified them as human, and 31.1% were unsure. These results suggest that faculty reviewers could not reliably distinguish AI-generated from human-written items.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Suitability for Exam Use\u003c/h2\u003e \u003cp\u003eReviewers also indicated whether each question was suitable for inclusion in a periodontology exam (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Among the human-written items, 55.7% were considered suitable for inclusion in a periodontology exam, 23.2% were not suitable, and 21.1% were considered acceptable only with modifications. By comparison, 84.1% of AI-generated items were judged suitable, 6.8% were not suitable, and 9.1% might be usable with modification. Thus, faculty were far more likely to recommend AI questions for direct examination use.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Composite Quality Scores\u003c/h2\u003e \u003cp\u003eThe composite score, derived by summing participant responses to the six recoded criteria (range 0\u0026ndash;24), provided an overall measure of item quality (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Human-written questions had a mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD composite score of 18.3\u0026thinsp;\u0026plusmn;\u0026thinsp;5.1 with a median of 18.0 and an interquartile range (IQR) of 6.0. AI-generated items achieved a higher mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD score of 20.7\u0026thinsp;\u0026plusmn;\u0026thinsp;4.9 with a median of 23.0 and the same IQR of 6.0. A generalized linear mixed model treating the question as a random effect and the source as a fixed effect indicated that AI generation increased the composite score by 2.38 points (standard error 0.54; t\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;4.43; p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e3.6 Criterion‑Level \u0026ldquo;Excellent\u0026rdquo; Ratings\u003c/h2\u003e \u003cp\u003eBecause the INBDE rubric was collapsed into three categories (poor\u0026thinsp;=\u0026thinsp;1, moderate\u0026thinsp;=\u0026thinsp;3, excellent\u0026thinsp;=\u0026thinsp;4), we compared the proportion of responses rated \u0026ldquo;excellent\u0026rdquo; for each quality criterion (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Across all six criteria, ChatGPT‑generated questions were more likely to receive top ratings, with the largest absolute advantage in clarity (\u0026asymp;\u0026thinsp;29 percentage points) and the smallest in fairness (\u0026asymp;\u0026thinsp;5 points). These findings indicate that AI items were frequently rated in the highest category across multiple criteria in this sample.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eProportion of \u0026ldquo;Excellent\u0026rdquo; Ratings Across INBDE Quality Criteria for Human- and AI-Generated Questions.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCriterion\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHuman \u0026ldquo;excellent\u0026rdquo; (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT \u0026ldquo;excellent\u0026rdquo; (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDifference (percentage points) *\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQuestion clarity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e34.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e62.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;28.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eContent accuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e44.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e64.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;20.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQuality of distractors\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e40.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e63.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;22.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFairness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e71.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;5.4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlignment with curriculum\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e58.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e72.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;13.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGrammar \u0026amp; language\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e46.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e68.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e+\u0026thinsp;22.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e*Percentage-point difference relative to the human-generated group, which was considered the reference.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. DISCUSSION","content":"\u003cp\u003eIn this study we evaluated the quality of AI-generated MCQs in Periodontology, comparing items created by ChatGPT-4o with those written by faculty, using the INBDE rubric as a standardized framework. Our findings revealed that AI-generated questions had a significantly higher composite quality score and, descriptively, received higher ratings across the six evaluated domains, and they were also more frequently deemed suitable for inclusion in summative examinations. Reviewers were also largely unable to determine whether the questions were produced by AI or by human authors. Collectively, these results suggest that ChatGPT-4o can generate high-quality, well-structured assessment items that are indistinguishable from, and sometimes superior to, traditional faculty-written questions. One possible explanation is that large language models can adhere strictly to explicit instructions, formatting rules, and item-writing guidelines, producing questions that are more consistently structured than those written by humans, who may vary in approach or inadvertently deviate from recommended standards.\u003c/p\u003e \u003cp\u003eThe significantly higher composite score for AI-generated items aligns with findings from several recent evaluations of LLMs in health-professions assessment. Artsi et al. (2024), for example, compared ChatGPT-4, ChatGPT-3.5, and Med-PaLM 2 in generating USMLE-style items and found that all three models produced coherent, blueprint-aligned questions, with ChatGPT-4 yielding the highest editorial quality. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] In emergency medicine, Law et al. (2023) reported that ChatGPT-4o generated MCQs were grammatically strong but significantly easier and less discriminative than expert-written items. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] Ahmed et al. (2025) showed that ChatGPT and Bard could produce linguistically acceptable dental-caries MCQs, although many required refinements for depth, and Dave et al. (2025) found that AI-platforms as ChatGPT-4o, Grok 2, and Gemini, generated coherent undergraduate dentistry items but sometimes included double negatives, overly long stems, or factual errors. [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] In Periodontology, Ma et al. (2025) demonstrated that ChatGPT-4-generated exams produced higher student scores but poorer discrimination between high- and low-performers, suggesting insufficient cognitive challenge. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] In contrast to these mixed findings, our results show that ChatGPT-4o outperformed human-written items across all six INBDE domains and was more often judged suitable for exam use. However, consistent with the broader literature, we also did not evaluate psychometric performance, so this observed higher editorial quality cannot yet be interpreted as evidence of greater cognitive rigor.\u003c/p\u003e \u003cp\u003eIn our study, AI-generated items had higher proportions of top ratings for clarity, content accuracy, and curricular alignment in descriptive comparisons. Similar findings have been reported elsewhere. For example, Artsi et al. (2024) showed that ChatGPT-4 produced USMLE-style items with fewer structural flaws than human authors, and Cheung et al. (2023) found that ChatGPT-based urology questions were often more grammatically polished. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] Dental studies, including Dave et al. (2025), likewise noted that models such as ChatGPT-4o could generate items that read more clearly than faculty drafts. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] However, these same investigations also highlighted that highly polished items are not always conceptually accurate or fully on-topic. Thus, while structured prompts and context scaffolding likely contributed to the strong performance we observed, expert review remains essential to ensure that fidelity to the intended curriculum.\u003c/p\u003e \u003cp\u003eThe superior ratings likely reflect the model\u0026rsquo;s ability to adhere to established principles for MCQs, which demand a clear stem, a focused lead-in, and plausible distractors of similar length and grammatical structure. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] Common pitfalls include grammatical or logical cues, use of extreme terms (\u0026ldquo;always\u0026rdquo;, \u0026ldquo;never\u0026rdquo;), and distractors that are obviously wrong. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] In our study, the AI consistently avoided these flaws and provided balanced distractors, demonstrating a level of rule-following and consistency that is difficult for human authors to maintain over large item sets. This ability to consistently follow item-writing guidelines may contribute to the higher composite quality ratings observed for AI-generated items. However, high-quality questions also require alignment with learning outcomes and validity evidence, such as item difficulty and discrimination indices. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] Because we did not administer the questions to students, psychometric qualities remain unknown. Post-hoc analysis of item statistics and internal structure will be needed to confirm whether the higher expert ratings translate into reliable discrimination among learners. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eAn additional finding of this study was that reviewers could not reliably determine whether an item had been written by AI or by a human. This suggests that AI-generated questions now reach a level of fluency and structure that closely resembles expert authorship [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. From an academic-integrity standpoint, this indistinguishability highlights the need for clear safeguards, because perceived authorship can no longer serve as an indicator of quality. Rigorous, rubric-based review should therefore remain standard practice for all items, regardless of their source. The fact that AI-generated questions were more often judged immediately suitable for exam use also suggests that this technology may serve as a useful supplement to faculty-driven assessment development.\u003c/p\u003e \u003cp\u003eThe implications for dental education are significant. Integrating AI into assessment workflows could help address one of the most persistent challenges in curriculum design: the shortage of time and faculty resources for generating new, high-quality test items [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Timing advantages have been clearly demonstrated; for example, Cheung et al. (2023) reported that their ChatGPT-based system generated urology MCQs in a fraction of the time required for faculty authors, while maintaining acceptable editorial quality, often within seconds. [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] With LLMs capable of producing questions that meet or exceed linguistic standards, educators may shift their focus toward validation, cognitive-level enhancement, and psychometric calibration rather than first-draft item creation [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. This could enable the development of larger, more diverse, and continuously updated question banks.\u003c/p\u003e \u003cp\u003eOur study\u0026rsquo;s design offered several notable strengths. It featured a fully blinded, head-to-head comparison that applied all six INBDE quality criteria. Fourteen evaluators participated, drawn from both full and part-time faculty, all of whom are American Board certified by the American Academy of Periodontology, which enhanced ecological validity. The study also reported both criterion level and composite outcomes to provide a more nuanced picture of item performance.\u003c/p\u003e \u003cp\u003eDespite these encouraging results, the study has several limitations. The work was conducted within a single institution and limited to the discipline of Periodontology, which may restrict the generalizability of the findings to other educational contexts. Additionally, the outcomes were based exclusively on expert ratings; we did not collect psychometric data such as item difficulty or student performance. Finally, the results are specific to a particular large language model, ChatGPT-4o, and the prompting approach used, both of which may influence item quality.\u003c/p\u003e \u003cp\u003eSeveral authors have cautioned against relying too heavily on AI for summative assessments without strong faculty involvement. Norcini et al. (2018) note that high-stakes exams still require careful blueprinting and psychometric review after administration, which are processes that remain fundamentally human-driven. [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] Others have highlighted the risk of construct underrepresentation, where AI-generated items may focus mainly on lower-level recall unless they are intentionally designed to assess higher-order thinking. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] These concerns are consistent with earlier findings showing that AI-generated questions, although polished, may have lower discriminative ability if used without iterative refinement.\u003c/p\u003e \u003cp\u003eFrom an implementation standpoint, best-practice frameworks increasingly advocate for a hybrid model in which AI serves as an initial drafting tool, followed by structured expert review and psychometric validation. [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] Within such a model, AI-generated items may enhance efficiency without compromising assessment integrity. Our findings support this approach, suggesting that while AI can substantially improve the efficiency and consistency of item development, its optimal use lies in augmentation rather than replacement of faculty expertise. Establishing clear governance structures, reviewer training, and documentation of AI involvement will be essential as these tools become more deeply integrated into dental education.\u003c/p\u003e \u003cp\u003eFuture research should include larger, multicenter samples across multiple specialties, embed AI-generated items in live examinations to capture psychometric properties, and compare different LLMs or retrieval-augmented architectures. Studies should also assess the time and cost savings of AI-assisted item generation, evaluate learner acceptance, and explore whether fine-tuned models further improve quality and cognitive level.\u003c/p\u003e"},{"header":"5. CONCLUSIONS","content":"\u003cp\u003eThis study demonstrates that ChatGPT-4o can generate periodontology MCQs that achieved a significantly higher composite quality score in this blinded evaluation. Reviewers were largely unable to distinguish AI-generated from human-authored questions, and most AI items were deemed suitable for direct examination use. These findings suggest that LLMs can serve as valuable adjuncts in creating dental assessment materials, helping educators expand and refresh question banks with remarkable efficiency.\u003c/p\u003e \u003cp\u003eHowever, while the editorial and structural quality of AI-generated questions was strong, this does not necessarily translate into higher cognitive rigor or validity in discriminating learner performance. Expert oversight, therefore, remains essential to verify factual accuracy, capture clinical nuances, and maintain appropriate cognitive demand. Rather than replacing human judgment, AI should be viewed as a supportive tool that can streamline the drafting process while faculty remain central to content validation and educational integrity.\u003c/p\u003e \u003cp\u003eFuture work should integrate AI-generated questions into live assessments to examine psychometric performance and explore strategies to optimize collaboration between educators and generative models in dental education.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study protocol received approval from the Tufts Health Sciences Institutional Review Board (STUDY00005769) and was classified as minimal‑risk research.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for Publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors in this research provide consent for publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of Data and Materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets supporting the conclusions of this article are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors in this research declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBA designed the study, acquired \u0026nbsp;the data, and wrote the manuscript. LV acquired the data and revised the manuscript. SJ interpreted the data and performed the data analysis. KS wrote the manuscript. NK critically revised the manuscript. NJ critically revised the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eArtsi M, Karras BT, Palanica A. Evaluating the quality of AI-generated USMLE-style exam questions across ChatGPT-4, GPT-3.5, and Med-PaLM 2. BMJ Open. 2024;14:e082559. doi:10.1136/bmjopen-2023-082559.\u003c/li\u003e\n \u003cli\u003eKim JK, Kim SJ, Park JY, Ahn HS. Evaluating ChatGPT-4 in generating and assessing multiple-choice questions in urology. Int J Med Educ. 2025;16:23\u0026ndash;29. doi:10.5116/ijme.6623.faf2\u003c/li\u003e\n \u003cli\u003eAbouzeid E, Wassef R, Jawwad A, Harris P. Chatbots\u0026rsquo; role in generating single best answer questions for undergraduate medical student assessment: comparative analysis. JMIR Med Educ. 2025;11:e69521. doi:10.2196/69521\u003c/li\u003e\n \u003cli\u003eMa X, Pan W, Yu X. Evaluating AI-generated examination papers in periodontology: a comparative study with human-designed counterparts. BMC Med Educ. 2025;25:1099. doi:10.1186/s12909-025-07706-6\u003c/li\u003e\n \u003cli\u003eLaw AKK, Law KIW, Law MF, Lee KKC. Evaluation of multiple-choice questions generated by ChatGPT in emergency medicine: comparative cross-sectional study. JMIR Med Educ. 2023;9:e49174. doi:10.2196/49174\u003c/li\u003e\n \u003cli\u003eCheung BHH, Sim JJL, Lam ST, Wong RHL, Lai VWY. Evaluating the effectiveness and validity of artificial intelligence\u0026ndash;generated urology multiple-choice questions. Adv Med Educ Pract. 2023;14:1227\u0026ndash;1235. doi:10.2147/AMEP.S423138\u003c/li\u003e\n \u003cli\u003eKıyak YS, Coşkun AK, Kaymak Ş, Coşkun \u0026Ouml;, Budakoğlu İİ. Can ChatGPT generate surgical multiple-choice questions comparable to those written by a surgeon? Proc (Bayl Univ Med Cent). 2024;38(1):48\u0026ndash;52. doi:10.1080/08998280.2024.2418752\u003c/li\u003e\n \u003cli\u003eKlang E, Perelman V, Avnon T, Klang S, Halpern P. Utility of ChatGPT in multiple choice question writing for medical students. BMC Med Educ. 2023;23:908. doi:10.1186/s12909-023-04669-1\u003c/li\u003e\n \u003cli\u003eEmekli E, Yıldırım S, Aydın E. Can ChatGPT generate appropriate radiology multiple-choice questions for medical students? A pilot study. Clin Imaging. 2024;98:44\u0026ndash;49. doi:10.1016/j.clinimag.2023.11.005\u003c/li\u003e\n \u003cli\u003eAhmed WM, Azhari AA, Alfaraj A, Alhamadani A, Zhang M, Lu CT. The quality of AI-generated dental caries multiple-choice questions: a comparative analysis of ChatGPT and Google Bard language models. Heliyon. 2024;10:e28198. doi:10.1016/j.heliyon.2024.e28198\u003c/li\u003e\n \u003cli\u003eDave M, Tattar R, Alafaleg R, Barry S, Ariyaratnam S, Roudsari RV, et al. Performance of large language models (ChatGPT-4o, Grok-2, and Gemini) in UK dentistry and dental hygiene and therapy assessments. Br Dent J. 2025. doi:10.1038/s41415-025-8383-2\u003c/li\u003e\n \u003cli\u003eAhmed A, Kerr E, O\u0026rsquo;Malley A. Quality assurance and validity of AI-generated single best answer questions. BMC Med Educ. 2025;25:300. doi:10.1186/s12909-025-06881-w\u003c/li\u003e\n \u003cli\u003eGottlieb M, Bailitz J, Fix M, Shappell E. Educator\u0026rsquo;s blueprint: a how-to guide for developing high-quality multiple-choice questions. AEM Educ Train. 2023;7(1):e10836. doi:10.1002/aet2.10836\u003c/li\u003e\n \u003cli\u003eNaidoo M. The pearls and pitfalls of setting high-quality multiple choice questions for clinical medicine. S Afr Fam Pract. 2023;65(1):e1\u0026ndash;e4. doi:10.4102/safp.v65i1.5726\u003c/li\u003e\n \u003cli\u003eDunlap CA, Lai G. Enhancing dental education through AI: revolutionizing exam creation and grading [Internet]. American Association of Endodontists; 2024 Dec 20 [cited 2026 Jan 2]. Available from: https://www.aae.org/specialty/enhancing-dental-education-through-ai-revolutionizing-exam-creation-and-grading/\u003c/li\u003e\n \u003cli\u003eNorcini J, Anderson MB, Bollela V, et al. 2018 consensus framework for good assessment. \u003cstrong\u003eMed Teach\u003c/strong\u003e. 2018;40(11):1102\u0026ndash;1109.\u003c/li\u003e\n \u003cli\u003eBloom BS, Engelhart MD, Furst EJ, Hill WH, Krathwohl DR. \u003cstrong\u003eTaxonomy of educational objectives: the classification of educational goals\u003c/strong\u003e. New York (NY): Longmans; 1956.\u003c/li\u003e\n \u003cli\u003eSchuwirth LWT, van der Vleuten CPM. Programmatic assessment: from assessment of learning to assessment for learning. \u003cstrong\u003eMed Teach\u003c/strong\u003e. 2011;33(6):478\u0026ndash;485.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Tufts University School of Dental Medicine","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"artificial intelligence, periodontology, dental education, multiple-choice questions, educational assessment","lastPublishedDoi":"10.21203/rs.3.rs-8614641/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8614641/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e This study evaluated the quality of multiple-choice questions (MCQs) generated by ChatGPT-4o compared with faculty written items in periodontology using the Integrated National Board Dental Examination (INBDE) rubric.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e Thirty MCQs were assessed in a blinded cross-sectional comparison at Tufts University School of Dental Medicine. 15 questions were generated by ChatGPT-4o based on course objectives and INBDE guidelines, and 15 were randomly selected from the departmental exam bank. Fourteen periodontology faculty members rated each item on six INBDE criteria including clarity, content accuracy, distractor quality, fairness, curricular alignment, and grammar using a five-point Likert scale ranging from poor (1) to excellent (5). Composite scores were analyzed using a generalized linear mixed model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e AI-generated items achieved significantly higher composite scores than human-written questions (20.7 ± 4.9 vs 18.3 ± 5.1; p \u0026lt; 0.001). In descriptive comparisons, AI items also received higher ratings across all six domains, particularly in clarity and grammar.\u003c/p\u003e\n\u003cp\u003eReviewers were unable to reliably identify the source of the items, and 84.1% of AI generated questions were judged suitable for exam use compared with 55.7% of faculty written items.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e ChatGPT-4o produced high-quality and well-structured MCQs, and reviewers frequently reported difficulty distinguishing their origin in this blinded assessment.\u003c/p\u003e\n\u003cp\u003eWhile these results highlight the potential value of AI-assisted assessment design, expert supervision remains essential to ensure accuracy, cognitive depth, and alignment with educational standards. AI should be a supportive tool that complements rather than replaces faculty expertise in question development.\u003c/p\u003e","manuscriptTitle":"Evaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-19 06:28:12","doi":"10.21203/rs.3.rs-8614641/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"344a2ae2-7f55-475d-ac13-0846c670f6ed","owner":[],"postedDate":"January 19th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":61216172,"name":"Dentistry"},{"id":61216173,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2026-01-19T06:28:13+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-19 06:28:12","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8614641","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8614641","identity":"rs-8614641","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00