Benchmarking Large Language Models on Persian Surgical Subspecialty Board Examinations: A Comparative Study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Benchmarking Large Language Models on Persian Surgical Subspecialty Board Examinations: A Comparative Study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash Shahab Sheikhalishahi, Farzad Rafiei, Seyed Masoud Hosseini, Alireza Haddadi, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8080315/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract This study evaluated the performance of three large language models, including ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash, on 532 Persian multiple-choice questions from the 2025 Iranian surgical subspecialty board examinations. Questions spanned five domains: Pediatric, Cardiovascular, Vascular and Endovascular, Thoracic, and Plastic & Reconstructive Surgery. Using standardized prompts, we assessed overall accuracy, variation across subspecialties and question types, and the effect of question length. ChatGPT-5 and Gemini 2.5 Flash achieved higher accuracy (73.3% and 73.9%) than ChatGPT-4o (68.2%). Agreement with the official key was substantial for Gemini 2.5 Flash (κ = 0.651) and ChatGPT-5 (κ = 0.642), and moderate to substantial for ChatGPT-4o (κ = 0.575). Model performance was stable across subspecialties, but all three showed lower accuracy on surgical technique questions compared with clinical scenarios or basic science items. Question length did not affect ChatGPT-5 or Gemini 2.5 Flash, while longer stems reduced ChatGPT-4o’s performance. These findings indicate that newer LLMs provide measurable improvements in surgical question answering, though persistent limitations in procedural reasoning suggest the need for careful integration and further multimodal development. Health sciences/Health care Health sciences/Medical research Surgery Artificial Intelligence (AI) ChatGPT Gemini Figures Figure 1 Figure 2 Figure 3 Introduction In recent years, Large Language Models (LLMs) such as GPT and Gemini have rapidly advanced in their capabilities for natural language understanding and generation. These models are increasingly explored as tools in medical education, clinical decision support, and assessment( 1 – 3 ). Their expanding role highlights the need to rigorously evaluate their reliability in high-stakes domains such as surgery, where reasoning and decision-making are complex. Early investigations have shown that LLMs can pass portions of medical licensing exams and outperform earlier AI methods in many medical question domains( 4 – 7 ). These findings suggest that LLMs are not only capable of memorizing factual knowledge but may also demonstrate reasoning skills relevant to clinical practice. Within surgery, the performance of GPT-4 and GPT-3.5 in answering the Korean general surgery board examination has been evaluated. GPT-4 demonstrated an improvement over the earlier version, achieving an accuracy of 76.4% compared to 46.8%( 8 ). To our knowledge, no prior work has compared LLMs on a surgical subspecialty multiple-choice question (MCQ) set in Persian. In this study, we thus assess and contrast the accuracy of three leading LLMs on 532 Persian surgical subspecialty MCQs, analyzing performance across subspecialties, question types, and question length. As newer model versions emerge rapidly, direct head-to-head comparison (ChatGPT-4o vs ChatGPT-5 vs Gemini 2.5 Flash) shows whether architectural and training improvements translate into gains in surgical reasoning. The results provide insights into current strengths, systematic weakness patterns (especially in procedural reasoning), and implications for using LLMs as educational tools or assessment aids in surgical training. Methods Study Design and Setting We performed a cross-sectional benchmarking study to quantify and compare the answer accuracy of three LLMs - ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash - on Persian, high-stakes MCQs from surgical subspecialties. Five domains were evaluated: Pediatric Surgery, Cardiovascular Surgery, Vascular and Endovascular Surgery, Thoracic Surgery, and Plastic and Reconstructive Surgery. All model runs were executed in September 2025. No clinical sites, patient records, or institutional data were involved. Data Sources The corpus comprised the 2025 Iranian surgical subspecialty board examinations (Persian language). In total, 550 MCQs were collated across five domains: Pediatric Surgery, Cardiovascular Surgery, Vascular and Endovascular Surgery, Thoracic Surgery, and Plastic and Reconstructive Surgery. All items were obtained directly from the official Sanjesh Organization website after formal public release; the materials are publicly accessible, and no additional institutional approvals or permissions were required for research use. This collection constitutes an authoritative and comprehensive benchmark dataset for evaluating the performance of advanced LLMs, reflecting the rigor of national board standards and ensuring clinical relevance and educational robustness for model assessment. Eligibility criteria Inclusion. Persian-language MCQs from the 2025 surgical subspecialty examinations featuring four options with a single correct answer. Exclusion. Items requiring non-textual media (e.g., figures, radiographs, imaging panels, graphs) and items formally invalidated or withdrawn by the examination authority. Analytic sample. Applying the pre-specified exclusions yielded 532 analyzable questions (18 excluded). The distribution of analyzable items by subspecialty was: Pediatric Surgery (n = 98), Cardiovascular Surgery (n = 99), Vascular and Endovascular Surgery (n = 95), Thoracic Surgery (n = 91), and Plastic and Reconstructive Surgery (n = 149). Model Configuration and Prompting ChatGPT models (4o; 5) and Gemini (2.5 Flash) are general-purpose LLMs trained on broad text corpora and optimized for instruction following and reasoning. The specific builds evaluated were: ChatGPT-4o (March 2025), ChatGPT-5 (mid-2025), and Gemini 2.5 Flash (mid-2025). Each model was accessed via a single, fixed interface (API or web application held constant per model). To eliminate cross-item contamination, every question was submitted in an isolated, memory-free session, and separate accounts were used per model. All items were submitted in Persian, preserving the original stems and options. A uniform instruction was used for every model and item: “I will give you a multiple-choice question. You must: • State only the correct option. Answer format: • Correct option: …” Because the benchmark language is Persian, an equivalent Persian instruction was delivered identically to all three models. When a model produced text beyond the option label, the first valid option marker (A/B/C/D or 1/2/3/4) was recorded as the response. Outputs lacking a uniquely identifiable, valid option were deemed invalid and counted as incorrect. For length metrics, question length was defined as the word count of the stem only (answer options excluded). Outcomes and Measures Question typology. Each item was assigned to one of three types based on dominant cognitive demand, with brief illustrative prompts: Clinical scenario: case-based diagnostic/management reasoning (e.g., optimal biopsy approach for a suspected facial melanoma). Basic science: foundational concepts without extended vignettes (e.g., statements about mesh suture closure for anterior abdominal wall defects). Surgical technique: operative steps/instrumentation/technical decisions (e.g., exceptions regarding fenestration in Fontan surgery). Endpoints. Primary: item-level accuracy (correct/incorrect) per model. Secondary: (i) paired inter-model accuracy comparisons; (ii) subspecialty-specific accuracy; (iii) question-type–specific accuracy; (iv) association between stem word count and correctness. Evaluation procedures (summary). Accuracy vs. official key: item-level correctness computed against Sanjesh answer keys. Inter-model comparisons: item-matched, pairwise assessments to identify relative strengths/weaknesses. Subspecialty stratification: accuracy summarized across the five surgical subspecialties. Question-type analysis: performance compared across clinical scenario, basic science, and surgical technique items. Question-length analysis: mean stem word counts contrasted between correct and incorrect items within each model. Statistical Analysis Descriptive statistics were used to summarize the accuracy of each LLM across surgical subspecialties. To compare paired model performances on the same set of questions, McNemar’s test was applied with a significance threshold of α = 0.05. The Pearson Chi-square test was employed to examine associations between subspecialty and model accuracy, while Cohen’s Kappa was calculated to assess the level of agreement between each model and the answer key. An independent-samples t-test was performed to evaluate whether question word count differed between correctly and incorrectly answered items for each model. Finally, a Chi-square test of independence was used to determine whether model performance varied across question types (clinical scenario, basic, or surgical technique). Bias Mitigation and Reproducibility To minimize context and order effects, all models received items in the same fixed order, with fresh, memory-cleared sessions per item. Model interfaces and parameters were held constant across the evaluation. Ethics and Data Governance This study did not require formal ethical approval because it did not involve human participants, patient data, biological samples, or identifiable personal information. The dataset consisted exclusively of publicly available multiple-choice board examination questions from the Sanjesh Organization, and all responses were generated by artificial intelligence models. In accordance with journal guidelines, we confirm that the study adhered to all relevant academic and institutional standards for research integrity. These examination items were accessed after formal public release on the official Sanjesh website; they are publicly available and did not require additional permissions for research use. Results A total of 550 MCQ questions from various surgical subspecialties were selected. Eighteen questions were excluded due to containing images, graphs, or imaging data, leaving 532 questions to be evaluated by the models and subsequently analyzed. A McNemar test was performed to compare the performance of LLMs (N = 532, α = 0.05). Significant differences were found between ChatGPT-4o and ChatGPT-5 and between ChatGPT-4o and Gemini 2.5 Flash, while Gemini 2.5 Flash and ChatGPT-5 did not differ significantly (χ² = 0.038, p = 0.845) (Table 1). Table 1. McNemar Test Comparing Performance of LLMs LLM Chi-Square P-value ChatGPT-4o & ChatGPT-5 7.953 0.005 ChatGPT-4o & Gemini 2.5 Flash 7.250 0.007 Gemini 2.5 Flash & ChatGPT-5 0.038 0.845 N = 532, p< 0.05 For Gemini 2.5 Flash, the proportion of correct answers by subspecialty was as follows: 75 out of 98 questions in Pediatric Surgery (76.5%), 75 out of 99 in Cardiovascular Surgery (75.8%), 70 out of 95 in Vascular and Endovascular Surgery (73.7%), 70 out of 91 in Thoracic Surgery (76.9%), and 100 out of 149 in Plastic and Reconstructive Surgery (67.1%) (Figure 1). Overall, Gemini 2.5 Flash answered 393 out of 532 questions correctly (73.9%) (Table 2). For ChatGPT-5, the proportion of correct answers by subspecialty was: 75 out of 98 questions in Pediatric Surgery (76.5%), 75 out of 99 in Cardiovascular Surgery (75.8%), 70 out of 95 in Vascular and Endovascular Surgery (73.7%), 70 out of 91 in Thoracic Surgery (76.9%), and 100 out of 149 in Plastic and Reconstructive Surgery (67.1%). Overall, ChatGPT-5 answered 390 out of 532 questions correctly (73.3%). Table 2. Accuracy of LLMs on 532 Surgical MCQs LLM False True Frequency Percent Frequency Percent ChatGPT-4o 169 31.8 363 68.2 ChatGPT-5 142 26.7 390 73.3 Gemini 2.5 Flash 139 26.1 393 73.9 For ChatGPT-4o, the proportion of correct answers by subspecialty was: 65 out of 98 questions in Pediatric Surgery (66.3%), 68 out of 99 in Cardiovascular Surgery (68.7%), 63 out of 95 in Vascular and Endovascular Surgery (66.3%), 66 out of 91 in Thoracic Surgery (72.5%), and 101 out of 149 in Plastic and Reconstructive Surgery (67.8%). Overall, ChatGPT-4o answered 363 out of 532 questions correctly (68.2%). The Pearson Chi-square test indicated no statistically significant association between subspecialty and Gemini 2.5 Flash answers (p = 0.730), ChatGPT-5 answers (p = 0.360), or ChatGPT-4o answers (p = 0.891). Across 532 valid cases, Cohen’s Kappa was calculated to evaluate agreement between model responses and the answer key. Gemini demonstrated substantial agreement with the answer key, κ = 0.651, p < 0.001. GPT-5 also showed substantial agreement with the answer key, κ = 0.642, p < 0.001. In contrast, GPT-4 exhibited moderate to substantial agreement with the answer key, κ = 0.575, p < 0.001. An independent-samples t-test was conducted to investigate whether the word count of questions was related to model performance. For Gemini 2.5 Flash, questions answered correctly had a mean word count of 30.94 words (SD = 19.36), while those answered incorrectly had a mean word count of 33.78 words (SD = 21.76). This difference did not reach statistical significance (p = 0.175). For ChatGPT-5, correctly answered questions had a mean word count of 31.32 words (SD = 19.98), compared to 32.68 words (SD = 20.22) for incorrectly answered questions, with no statistically significant difference (p = 0.489). In contrast, for ChatGPT-4o, correctly answered questions had a mean word count of 30.46 words (SD = 19.56), whereas incorrectly answered questions averaged 34.29 words (SD = 20.84). This difference was statistically significant (p = 0.040). A chi-square test of independence was conducted to examine the relationship between question type and model performance. For Gemini 2.5 Flash, the distribution of correct and incorrect responses differed significantly across question types (p = 0.017). Specifically, Gemini 2.5 Flash achieved higher accuracy on clinical scenario questions (76.8% correct) and basic questions (77.4% correct), compared to surgical technique questions (64.4% correct). Similarly, for GPT-5, the distribution of correct and incorrect responses differed significantly across question types (p = 0.049), with higher accuracy on clinical scenario questions (75.5% correct) and basic questions (76.7% correct) than on surgical technique questions (65.2% correct). In contrast, for GPT-4, the distribution of correct and incorrect responses did not differ significantly across question types (p = 0.083). Accuracy was highest for clinical scenario questions (68.5% correct) and basic questions (73.6% correct), compared to surgical technique questions (61.4% correct). The performance of each model across different subspecialties and question types is presented in Figure 3. Discussion In this study, we systematically compared the performance of three LLMs, ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash, on 532 MCQs drawn from a broad range of surgical topics. Our principal findings are fourfold. First, ChatGPT-5 and Gemini 2.5 Flash achieved significantly higher accuracy than ChatGPT-4o (McNemar’s test). Second, model accuracy was consistent across surgical subspecialties. Third, question length showed no meaningful association with performance except for a modest decline in ChatGPT-4o. Finally, all models performed less accurately on surgical-technique questions compared with clinical scenario or basic factual items. The superior performance of ChatGPT-5 and Gemini 2.5 Flash suggests that ongoing architectural refinements and expanded training corpora yield measurable gains, even in highly specialized medical domains. The observed Cohen’s kappa values (Gemini 2.5 Flash κ = 0.651, ChatGPT-5 κ = 0.642, ChatGPT-4o κ = 0.575) fall within the substantial agreement range (κ ≈ 0.61–0.80), underscoring the robustness of these improvements. Similar upward trends have been documented in broader medical MCQ settings across successive LLM releases( 9 ). The lack of a significant McNemar difference between ChatGPT-5 and Gemini 2.5 Flash may reflect convergence in the capabilities of top-tier models. Performance consistency across subspecialties indicates that these models possess domain-general reasoning capacity rather than niche specialization. Question length had little impact on Gemini 2.5 Flash and ChatGPT-5, whereas ChatGPT-4o showed a small but significant decline with longer stems, consistent with prior reports of older LLMs being more sensitive to prompt complexity( 10 ). Despite these gains, all models showed persistent underperformance on surgical-technique questions (Gemini 2.5 Flash: 64.4% vs ~ 76% for clinical/basic items; ChatGPT-5: 65.2% vs ~ 75–77%). Such items often require procedural, spatial, or operative reasoning; skills less represented in text-based training data, whereas clinical scenarios may be more amenable to the pattern-recognition strategies in which LLMs excel. Prior studies have demonstrated that ChatGPT-4o and related models can meet or exceed passing thresholds on standardized medical examinations such as the USMLE( 11 ). However, few have focused specifically on surgical MCQs or stratified performance by question type. Our results extend this literature by highlighting inter-model differences within a surgical domain and by identifying procedural reasoning as a persistent challenge. Similar patterns of incremental improvement with narrowing margins have been observed in other medical specialties as LLMs mature( 12 ). Despite encouraging accuracy, the integration of LLMs into surgical education raises important ethical and equity concerns. Unsupervised use by trainees could undermine assessment integrity or foster overreliance on AI-generated explanations( 13 ). Furthermore, LLMs are trained on heterogeneous datasets that may contain cultural, gender, or regional biases, potentially influencing question interpretation or answer patterns in subtle ways( 14 ). Ongoing bias auditing, transparent reporting of training-data limitations, and clear institutional policies will be essential to ensure fair and responsible deployment. Rather than replacing human educators, LLMs may be most effective as collaborative partners that augment expert oversight. Educators can leverage these models for rapid question generation, preliminary grading, or formative feedback while retaining authority over content accuracy and educational context. Such hybrid systems could improve scalability and efficiency without displacing the nuanced judgment required for surgical training. Future work should explore optimal workflows for integrating LLMs’ outputs into instructor-led teaching and assessment. Widespread adoption will also depend on cost and infrastructure. Premium models such as ChatGPT-5 and Gemini 2.5 Flash often require paid subscriptions or institutional licenses, which may limit uptake in lower-resource settings. Computing demands, data privacy regulations, and internet connectivity further constrain global accessibility. Careful attention to pricing structures, data security, and open-source alternatives will be critical to ensure equitable integration. Finally, high MCQ accuracy does not guarantee clinically valid reasoning. LLMs may reach correct answers through statistical cues rather than genuine understanding. Evaluating confidence calibration, analyzing generated rationales, and comparing reasoning chains with expert logic will be vital for establishing trust. Strengths and Limitations Key strengths of this study include a large, diverse set of surgical MCQs and the application of rigorous statistical methods, including McNemar’s tests and Cohen’s kappa, to quantify model agreement. Stratification by question type and length adds a level of granularity rarely reported in prior LLM evaluations. Nevertheless, several limitations warrant consideration. MCQs reward pattern recognition and may not fully capture deeper clinical reasoning or operative judgment. Image-based questions were excluded, potentially underestimating the advantage of newer multimodal models. Only a single zero-shot prompt was used; alternative strategies might yield different results. The performance of surgical residents or trainees was not included, limiting the educational context. The analysis reflects a single time point and does not account for longitudinal stability or model drift. Accuracy assumes correctness of the provided key. Implications and Future Directions The high accuracy of ChatGPT-5 and Gemini 2.5 Flash suggests potential utility as formative assessment tools, question-answer tutors, or aids in MCQ generation for surgical education. However, their consistent deficits in surgical technique raise caution against using these models for procedural instruction without human oversight. Future research should examine hybrid assessment methods that include free-text, simulations, or image-based tasks, optimize prompts, and use chain-of-thought reasoning to improve procedural accuracy, incorporate multimodal inputs like operative images or videos as new models develop visual reasoning abilities, and compare results to human trainees to set meaningful performance benchmarks. Conclusion Newer LLMs, ChatGPT-5 and Gemini 2.5 Flash, significantly outperform ChatGPT-4o on a diverse set of surgical MCQs and achieve substantial agreement with expert answer keys. Their performance is consistent across surgical subspecialties and largely independent of question length, yet all models exhibit persistent deficits in surgical-technique reasoning. These findings support the cautious integration of advanced LLMs into surgical education and assessment while underscoring the need for multimodal approaches, refined prompting, and ongoing human oversight. Declarations Author Contributions Shahab Sheikhalishahi and Alireza Haddadi contributed to data collection. Farzad Rafiei, Shahab Sheikhalishahi, Alireza Haddadi, and Saina Sadeghipour contributed to the conceptualization, methodology, drafting of the original manuscript, and its review and editing. Farzad Rafiei performed the data analysis. Seyed Masoud Hosseini contributed to supervision, as well as the review and editing of the manuscript. All authors read and approved the final version of the manuscript. Competing Interest The authors declare no competing interests. Data availability The dataset of 2025 Iranian surgical medicine subspecialty board questions is publicly available from the official Sanjesh website (https://sanjeshp.ir). Any additional processed data generated and analyzed during the current study are available from the corresponding authors. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. References Aster A, Laupichler MC, Rockwell-Kollmann T, Masala G, Bala E, Raupach T. ChatGPT and Other Large Language Models in Medical Education — Scoping Literature Review. Medical Science Educator. 2025;35(1):555-67. Busch F, Hoffmann L, Rueger C, van Dijk EHC, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine. 2025;5(1):26. Vrdoljak J, Boban Z, Vilović M, Kumrić M, Božić J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare (Basel). 2025;13(6). Lin CR, Chen YJ, Tsai PA, Hsieh WY, Tsai SHL, Fu TS, et al. Multiple large language models versus clinical guidelines for postmenopausal osteoporosis: a comparative study of ChatGPT-3.5, ChatGPT-4.0, ChatGPT-4o, Google Gemini, Google Gemini Advanced, and Microsoft Copilot. Arch Osteoporos. 2025;20(1):120. Clark KR. Comparative Analysis of LLMs' Performance On a Practice Radiography Certification Exam. Radiol Technol. 2025;96(5):334-42. Khan AA, Yunus R, Sohail M, Rehman TA, Saeed S, Bu Y, et al. Artificial Intelligence for Anesthesiology Board-Style Examination Questions: Role of Large Language Models. J Cardiothorac Vasc Anesth. 2024;38(5):1251-9. Passby L, Jenko N, Wernham A. Performance of ChatGPT on Specialty Certificate Examination in Dermatology multiple-choice questions. Clin Exp Dermatol. 2024;49(7):722-7. Oh N, Choi GS, Lee WY. ChatGPT goes to the operating room: evaluating GPT-4 performance and its potential in surgical education and training in the era of large language models. Ann Surg Treat Res. 2023;104(5):269-73. Elzayyat M, Mohammad JN, Zaqout S. Assessing LLM-generated vs. expert-created clinical anatomy MCQs: a student perception-based comparative study in medical education. Med Educ Online. 2025;30(1):2554678. Lu Y, Aleta A, Du C, Shi L, Moreno Y. LLMs and generative agent-based models for complex systems research. Phys Life Rev. 2024;51:283-93. Chen Y, Huang X, Yang F, Lin H, Lin H, Zheng Z, et al. Performance of ChatGPT and Bard on the medical licensing examinations varies across different cultures: a comparison study. BMC Med Educ. 2024;24(1):1372. Yu E, Chu X, Zhang W, Meng X, Yang Y, Ji X, et al. Large Language Models in Medicine: Applications, Challenges, and Future Directions. Int J Med Sci. 2025;22(11):2792-801. Zhai C, Wibowo S, Li LD. The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic review. Smart Learning Environments. 2024;11(1):28. Schwartz R, Schwartz R, Vassilev A, Greene K, Perine L, Burt A, et al. Towards a standard for identifying and managing bias in artificial intelligence: US Department of Commerce, National Institute of Standards and Technology …; 2022. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 06 Mar, 2026 Reviews received at journal 06 Mar, 2026 Reviewers agreed at journal 02 Mar, 2026 Reviewers agreed at journal 02 Mar, 2026 Reviews received at journal 23 Feb, 2026 Reviewers agreed at journal 14 Feb, 2026 Reviewers invited by journal 13 Feb, 2026 Editor assigned by journal 09 Feb, 2026 Editor invited by journal 26 Nov, 2025 Submission checks completed at journal 19 Nov, 2025 First submitted to journal 15 Nov, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8080315","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":591601775,"identity":"38bad7a3-d0a7-4648-9bd8-4d242a16fb98","order_by":0,"name":"Shahab Sheikhalishahi","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Shahab","middleName":"","lastName":"Sheikhalishahi","suffix":""},{"id":591601776,"identity":"1d8790b8-2ce4-462a-b9b7-f6c7d8830ec6","order_by":1,"name":"Farzad Rafiei","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Farzad","middleName":"","lastName":"Rafiei","suffix":""},{"id":591601777,"identity":"2bfc864b-93b2-4339-adb7-158abeafe3aa","order_by":2,"name":"Seyed Masoud Hosseini","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Seyed","middleName":"Masoud","lastName":"Hosseini","suffix":""},{"id":591601778,"identity":"66944a62-ed8c-4b23-95d5-308998e625dd","order_by":3,"name":"Alireza Haddadi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABGElEQVRIiWNgGAWjYBACAyCSYGCTYOBjBvNtePhBVEIBEVrYIFrS5CQbQFoMCGphACEQOGxscAAqjguYsx/eeONHmYU8Gzvvw4c/atISN59fnfjhgQGDPL/YAaxaLHvSii17zkkYtjGzGxtIHLNJ3Hbj7WYJoMMMZ85OwO6wAzlmErxtEoxtzGxsEgZsaUAtZzeAtCQY3Mah5fwbM8m/bRL2QC3sPxL+HU7cPOPs5h94tdzIMZMG2pIIsoXhYBvQ+/y92/DbcuNZsbXMOYlkoBZmyca+NDmJG7zbLBIMJHD75Xzyxptvyups+/mPMX788Q0Ylf1nN9/8UWEjzy+NXQsWIAFWKUGschDgP0CK6lEwCkbBKBgBAACt11xfOZG1LwAAAABJRU5ErkJggg==","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":true,"prefix":"","firstName":"Alireza","middleName":"","lastName":"Haddadi","suffix":""},{"id":591601779,"identity":"d6caf74d-600a-438d-84d7-37ae714a0525","order_by":4,"name":"Saina Sadeghipour","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Saina","middleName":"","lastName":"Sadeghipour","suffix":""}],"badges":[],"createdAt":"2025-11-10 20:08:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8080315/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8080315/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102991813,"identity":"fd6f55d1-e5dc-464e-8980-0b5381a61b08","added_by":"auto","created_at":"2026-02-19 11:35:17","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":194618,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNumber of correct answers provided by each LLM across surgical subspecialties.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8080315/v1/974847b5bbf6b5244136feaa.jpg"},{"id":102991814,"identity":"01f19618-aa03-4edb-a0e9-6d354854f55a","added_by":"auto","created_at":"2026-02-19 11:35:17","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":114838,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNumber of correct answers by ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash across different question types\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8080315/v1/f7257514e8f60c2ab3eeba55.jpg"},{"id":102991815,"identity":"e1fce957-eb6a-4d86-b45c-87e9250aff6d","added_by":"auto","created_at":"2026-02-19 11:35:18","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":190615,"visible":true,"origin":"","legend":"\u003cp\u003eComparative Performance of LLMs in Answering Surgical MCQs Across Subspecialties\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8080315/v1/2b42024a7cc0bc435a8dcc85.jpg"},{"id":102991829,"identity":"09b208f2-b947-422d-968c-7a3ecd8571e8","added_by":"auto","created_at":"2026-02-19 11:35:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1121625,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8080315/v1/3088cefc-108b-4ccf-b6b3-4199aab660fd.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Benchmarking Large Language Models on Persian Surgical Subspecialty Board Examinations: A Comparative Study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn recent years, Large Language Models (LLMs) such as GPT and Gemini have rapidly advanced in their capabilities for natural language understanding and generation. These models are increasingly explored as tools in medical education, clinical decision support, and assessment(\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Their expanding role highlights the need to rigorously evaluate their reliability in high-stakes domains such as surgery, where reasoning and decision-making are complex.\u003c/p\u003e \u003cp\u003eEarly investigations have shown that LLMs can pass portions of medical licensing exams and outperform earlier AI methods in many medical question domains(\u003cspan additionalcitationids=\"CR5 CR6\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). These findings suggest that LLMs are not only capable of memorizing factual knowledge but may also demonstrate reasoning skills relevant to clinical practice. Within surgery, the performance of GPT-4 and GPT-3.5 in answering the Korean general surgery board examination has been evaluated. GPT-4 demonstrated an improvement over the earlier version, achieving an accuracy of 76.4% compared to 46.8%(\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). To our knowledge, no prior work has compared LLMs on a surgical subspecialty multiple-choice question (MCQ) set in Persian.\u003c/p\u003e \u003cp\u003eIn this study, we thus assess and contrast the accuracy of three leading LLMs on 532 Persian surgical subspecialty MCQs, analyzing performance across subspecialties, question types, and question length. As newer model versions emerge rapidly, direct head-to-head comparison (ChatGPT-4o vs ChatGPT-5 vs Gemini 2.5 Flash) shows whether architectural and training improvements translate into gains in surgical reasoning. The results provide insights into current strengths, systematic weakness patterns (especially in procedural reasoning), and implications for using LLMs as educational tools or assessment aids in surgical training.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eStudy Design and Setting\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe performed a cross-sectional benchmarking study to quantify and compare the answer accuracy of three LLMs - ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash - on Persian, high-stakes MCQs from surgical subspecialties. Five domains were evaluated: Pediatric Surgery, Cardiovascular Surgery, Vascular and Endovascular Surgery, Thoracic Surgery, and Plastic and Reconstructive Surgery. All model runs were executed in September 2025. No clinical sites, patient records, or institutional data were involved.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Sources\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe corpus comprised the 2025 Iranian surgical subspecialty board examinations (Persian language). In total, 550 MCQs were collated across five domains: Pediatric Surgery, Cardiovascular Surgery, Vascular and Endovascular Surgery, Thoracic Surgery, and Plastic and Reconstructive Surgery. All items were obtained directly from the official Sanjesh Organization website after formal public release; the materials are publicly accessible, and no additional institutional approvals or permissions were required for research use. This collection constitutes an authoritative and comprehensive benchmark dataset for evaluating the performance of advanced LLMs, reflecting the rigor of national board standards and ensuring clinical relevance and educational robustness for model assessment.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEligibility criteria\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eInclusion. Persian-language MCQs from the 2025 surgical subspecialty examinations featuring four options with a single correct answer.\u003c/p\u003e\n\u003cp\u003eExclusion. Items requiring non-textual media (e.g., figures, radiographs, imaging panels, graphs) and items formally invalidated or withdrawn by the examination authority.\u003c/p\u003e\n\u003cp\u003eAnalytic sample. Applying the pre-specified exclusions yielded 532 analyzable questions (18 excluded). The distribution of analyzable items by subspecialty was: Pediatric Surgery (n = 98), Cardiovascular Surgery (n = 99), Vascular and Endovascular Surgery (n = 95), Thoracic Surgery (n = 91), and Plastic and Reconstructive Surgery (n = 149).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Configuration and Prompting\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eChatGPT models (4o; 5) and Gemini (2.5 Flash) are general-purpose LLMs trained on broad text corpora and optimized for instruction following and reasoning. The specific builds evaluated were: ChatGPT-4o (March 2025), ChatGPT-5 (mid-2025), and Gemini 2.5 Flash (mid-2025). Each model was accessed via a single, fixed interface (API or web application held constant per model). To eliminate cross-item contamination, every question was submitted in an isolated, memory-free session, and separate accounts were used per model.\u003c/p\u003e\n\u003cp\u003eAll items were submitted in Persian, preserving the original stems and options. A uniform instruction was used for every model and item:\u003c/p\u003e\n\u003cp\u003e\u0026ldquo;I will give you a multiple-choice question.\u003cbr\u003e\u0026nbsp;You must:\u003cbr\u003e\u0026nbsp;\u0026bull; State only the correct option.\u003cbr\u003e\u0026nbsp;Answer format:\u003cbr\u003e\u0026nbsp;\u0026bull; Correct option: \u0026hellip;\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eBecause the benchmark language is Persian, an equivalent Persian instruction was delivered identically to all three models. When a model produced text beyond the option label, the first valid option marker (A/B/C/D or 1/2/3/4) was recorded as the response. Outputs lacking a uniquely identifiable, valid option were deemed invalid and counted as incorrect. For length metrics, question length was defined as the word count of the stem only (answer options excluded).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOutcomes and Measures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eQuestion typology. Each item was assigned to one of three types based on dominant cognitive demand, with brief illustrative prompts:\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003eClinical scenario: case-based diagnostic/management reasoning (e.g., optimal biopsy approach for a suspected facial melanoma).\u003c/li\u003e\n \u003cli\u003eBasic science: foundational concepts without extended vignettes (e.g., statements about mesh suture closure for anterior abdominal wall defects).\u003c/li\u003e\n \u003cli\u003eSurgical technique: operative steps/instrumentation/technical decisions (e.g., exceptions regarding fenestration in Fontan surgery).\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eEndpoints.\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003ePrimary: item-level accuracy (correct/incorrect) per model.\u003c/li\u003e\n \u003cli\u003eSecondary: (i) paired inter-model accuracy comparisons; (ii) subspecialty-specific accuracy; (iii) question-type\u0026ndash;specific accuracy; (iv) association between stem word count and correctness.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eEvaluation procedures (summary).\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003eAccuracy vs. official key: item-level correctness computed against Sanjesh answer keys.\u003c/li\u003e\n \u003cli\u003eInter-model comparisons: item-matched, pairwise assessments to identify relative strengths/weaknesses.\u003c/li\u003e\n \u003cli\u003eSubspecialty stratification: accuracy summarized across the five surgical subspecialties.\u003c/li\u003e\n \u003cli\u003eQuestion-type analysis: performance compared across clinical scenario, basic science, and surgical technique items.\u003c/li\u003e\n \u003cli\u003eQuestion-length analysis: mean stem word counts contrasted between correct and incorrect items within each model.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDescriptive statistics were used to summarize the accuracy of each LLM across surgical subspecialties. To compare paired model performances on the same set of questions, McNemar\u0026rsquo;s test was applied with a significance threshold of \u0026alpha; = 0.05. The Pearson Chi-square test was employed to examine associations between subspecialty and model accuracy, while Cohen\u0026rsquo;s Kappa was calculated to assess the level of agreement between each model and the answer key. An independent-samples t-test was performed to evaluate whether question word count differed between correctly and incorrectly answered items for each model. Finally, a Chi-square test of independence was used to determine whether model performance varied across question types (clinical scenario, basic, or surgical technique).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBias Mitigation and Reproducibility\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo minimize context and order effects, all models received items in the same fixed order, with fresh, memory-cleared sessions per item. Model interfaces and parameters were held constant across the evaluation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics and Data Governance\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study did not require formal ethical approval because it did not involve human participants, patient data, biological samples, or identifiable personal information. The dataset consisted exclusively of publicly available multiple-choice board examination questions from the Sanjesh Organization, and all responses were generated by artificial intelligence models. In accordance with journal guidelines, we confirm that the study adhered to all relevant academic and institutional standards for research integrity.\u003c/p\u003e\n\u003cp\u003eThese examination items were accessed after formal public release on the official Sanjesh website; they are publicly available and did not require additional permissions for research use.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eA total of 550 MCQ questions from various surgical subspecialties were selected. Eighteen questions were excluded due to containing images, graphs, or imaging data, leaving 532 questions to be evaluated by the models and subsequently analyzed. A McNemar test was performed to compare the performance of LLMs (N = 532, \u0026alpha; = 0.05). Significant differences were found between ChatGPT-4o and ChatGPT-5 and between ChatGPT-4o and Gemini 2.5 Flash, while Gemini 2.5 Flash and ChatGPT-5 did not differ significantly (\u0026chi;\u0026sup2; = 0.038, p = 0.845) (Table 1).\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" align=\"\" width=\"398\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" style=\"width: 398px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTable 1. McNemar Test Comparing Performance of LLMs\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 206px;\"\u003e\n \u003cp\u003eLLM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 108px;\"\u003e\n \u003cp\u003eChi-Square\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003eP-value\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 206px;\"\u003e\n \u003cp\u003eChatGPT-4o \u0026amp; ChatGPT-5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 108px;\"\u003e\n \u003cp\u003e7.953\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.005\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 206px;\"\u003e\n \u003cp\u003eChatGPT-4o \u0026amp; Gemini 2.5 Flash\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 108px;\"\u003e\n \u003cp\u003e7.250\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.007\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 206px;\"\u003e\n \u003cp\u003eGemini 2.5 Flash \u0026amp; ChatGPT-5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 108px;\"\u003e\n \u003cp\u003e0.038\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 84px;\"\u003e\n \u003cp\u003e0.845\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" style=\"width: 398px;\"\u003e\n \u003cp\u003eN = 532, p\u0026lt; 0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFor Gemini 2.5 Flash, the proportion of correct answers by subspecialty was as follows: 75 out of 98 questions in Pediatric Surgery (76.5%), 75 out of 99 in Cardiovascular Surgery (75.8%), 70 out of 95 in Vascular and Endovascular Surgery (73.7%), 70 out of 91 in Thoracic Surgery (76.9%), and 100 out of 149 in Plastic and Reconstructive Surgery (67.1%) (Figure 1). Overall, Gemini 2.5 Flash answered 393 out of 532 questions correctly (73.9%) (Table 2). For ChatGPT-5, the proportion of correct answers by subspecialty was: 75 out of 98 questions in Pediatric Surgery (76.5%), 75 out of 99 in Cardiovascular Surgery (75.8%), 70 out of 95 in Vascular and Endovascular Surgery (73.7%), 70 out of 91 in Thoracic Surgery (76.9%), and 100 out of 149 in Plastic and Reconstructive Surgery (67.1%). Overall, ChatGPT-5 answered 390 out of 532 questions correctly (73.3%). \u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" align=\"\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"5\" valign=\"top\" style=\"width: 399px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTable 2. Accuracy of LLMs on 532 Surgical MCQs\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"2\" valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/p\u003e\n \u003cp\u003eLLM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 142px;\"\u003e\n \u003cp\u003eFalse\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 137px;\"\u003e\n \u003cp\u003eTrue\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eFrequency\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 65px;\"\u003e\n \u003cp\u003ePercent\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003eFrequency\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003ePercent\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eChatGPT-4o\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e169\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 65px;\"\u003e\n \u003cp\u003e31.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e363\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e68.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eChatGPT-5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e142\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 65px;\"\u003e\n \u003cp\u003e26.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e390\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e73.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 120px;\"\u003e\n \u003cp\u003eGemini 2.5 Flash\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e139\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 65px;\"\u003e\n \u003cp\u003e26.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 77px;\"\u003e\n \u003cp\u003e393\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 61px;\"\u003e\n \u003cp\u003e73.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFor ChatGPT-4o, the proportion of correct answers by subspecialty was: 65 out of 98 questions in Pediatric Surgery (66.3%), 68 out of 99 in Cardiovascular Surgery (68.7%), 63 out of 95 in Vascular and Endovascular Surgery (66.3%), 66 out of 91 in Thoracic Surgery (72.5%), and 101 out of 149 in Plastic and Reconstructive Surgery (67.8%). Overall, ChatGPT-4o answered 363 out of 532 questions correctly (68.2%). The Pearson Chi-square test indicated no statistically significant association between subspecialty and Gemini 2.5 Flash answers (p = 0.730), ChatGPT-5 answers (p = 0.360), or ChatGPT-4o answers (p = 0.891). Across 532 valid cases, Cohen\u0026rsquo;s Kappa was calculated to evaluate agreement between model responses and the answer key. Gemini demonstrated substantial agreement with the answer key, \u0026kappa; = 0.651, p \u0026lt; 0.001. GPT-5 also showed substantial agreement with the answer key, \u0026kappa; = 0.642, p \u0026lt; 0.001. In contrast, GPT-4 exhibited moderate to substantial agreement with the answer key, \u0026kappa; = 0.575, p \u0026lt; 0.001.\u003c/p\u003e\n\u003cp\u003eAn independent-samples t-test was conducted to investigate whether the word count of questions was related to model performance. For Gemini 2.5 Flash, questions answered correctly had a mean word count of 30.94 words (SD = 19.36), while those answered incorrectly had a mean word count of 33.78 words (SD = 21.76). This difference did not reach statistical significance (p = 0.175). For ChatGPT-5, correctly answered questions had a mean word count of 31.32 words (SD = 19.98), compared to 32.68 words (SD = 20.22) for incorrectly answered questions, with no statistically significant difference (p = 0.489). In contrast, for ChatGPT-4o, correctly answered questions had a mean word count of 30.46 words (SD = 19.56), whereas incorrectly answered questions averaged 34.29 words (SD = 20.84). This difference was statistically significant (p = 0.040).\u003c/p\u003e\n\u003cp\u003eA chi-square test of independence was conducted to examine the relationship between question type and model performance. For Gemini 2.5 Flash, the distribution of correct and incorrect responses differed significantly across question types (p = 0.017). Specifically, Gemini 2.5 Flash achieved higher accuracy on clinical scenario questions (76.8% correct) and basic questions (77.4% correct), compared to surgical technique questions (64.4% correct). Similarly, for GPT-5, the distribution of correct and incorrect responses differed significantly across question types (p = 0.049), with higher accuracy on clinical scenario questions (75.5% correct) and basic questions (76.7% correct) than on surgical technique questions (65.2% correct). In contrast, for GPT-4, the distribution of correct and incorrect responses did not differ significantly across question types (p = 0.083). Accuracy was highest for clinical scenario questions (68.5% correct) and basic questions (73.6% correct), compared to surgical technique questions (61.4% correct). The performance of each model across different subspecialties and question types is presented in Figure 3.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we systematically compared the performance of three LLMs, ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash, on 532 MCQs drawn from a broad range of surgical topics. Our principal findings are fourfold. First, ChatGPT-5 and Gemini 2.5 Flash achieved significantly higher accuracy than ChatGPT-4o (McNemar\u0026rsquo;s test). Second, model accuracy was consistent across surgical subspecialties. Third, question length showed no meaningful association with performance except for a modest decline in ChatGPT-4o. Finally, all models performed less accurately on surgical-technique questions compared with clinical scenario or basic factual items.\u003c/p\u003e \u003cp\u003eThe superior performance of ChatGPT-5 and Gemini 2.5 Flash suggests that ongoing architectural refinements and expanded training corpora yield measurable gains, even in highly specialized medical domains. The observed Cohen\u0026rsquo;s kappa values (Gemini 2.5 Flash κ\u0026thinsp;=\u0026thinsp;0.651, ChatGPT-5 κ\u0026thinsp;=\u0026thinsp;0.642, ChatGPT-4o κ\u0026thinsp;=\u0026thinsp;0.575) fall within the \u003cem\u003esubstantial agreement\u003c/em\u003e range (κ\u0026thinsp;\u0026asymp;\u0026thinsp;0.61\u0026ndash;0.80), underscoring the robustness of these improvements. Similar upward trends have been documented in broader medical MCQ settings across successive LLM releases(\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). The lack of a significant McNemar difference between ChatGPT-5 and Gemini 2.5 Flash may reflect convergence in the capabilities of top-tier models.\u003c/p\u003e \u003cp\u003ePerformance consistency across subspecialties indicates that these models possess domain-general reasoning capacity rather than niche specialization. Question length had little impact on Gemini 2.5 Flash and ChatGPT-5, whereas ChatGPT-4o showed a small but significant decline with longer stems, consistent with prior reports of older LLMs being more sensitive to prompt complexity(\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDespite these gains, all models showed persistent underperformance on surgical-technique questions (Gemini 2.5 Flash: 64.4% vs\u0026thinsp;~\u0026thinsp;76% for clinical/basic items; ChatGPT-5: 65.2% vs\u0026thinsp;~\u0026thinsp;75\u0026ndash;77%). Such items often require procedural, spatial, or operative reasoning; skills less represented in text-based training data, whereas clinical scenarios may be more amenable to the pattern-recognition strategies in which LLMs excel.\u003c/p\u003e \u003cp\u003ePrior studies have demonstrated that ChatGPT-4o and related models can meet or exceed passing thresholds on standardized medical examinations such as the USMLE(\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). However, few have focused specifically on surgical MCQs or stratified performance by question type. Our results extend this literature by highlighting inter-model differences within a surgical domain and by identifying procedural reasoning as a persistent challenge. Similar patterns of incremental improvement with narrowing margins have been observed in other medical specialties as LLMs mature(\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDespite encouraging accuracy, the integration of LLMs into surgical education raises important ethical and equity concerns. Unsupervised use by trainees could undermine assessment integrity or foster overreliance on AI-generated explanations(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Furthermore, LLMs are trained on heterogeneous datasets that may contain cultural, gender, or regional biases, potentially influencing question interpretation or answer patterns in subtle ways(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). Ongoing bias auditing, transparent reporting of training-data limitations, and clear institutional policies will be essential to ensure fair and responsible deployment.\u003c/p\u003e \u003cp\u003eRather than replacing human educators, LLMs may be most effective as collaborative partners that augment expert oversight. Educators can leverage these models for rapid question generation, preliminary grading, or formative feedback while retaining authority over content accuracy and educational context. Such hybrid systems could improve scalability and efficiency without displacing the nuanced judgment required for surgical training. Future work should explore optimal workflows for integrating LLMs\u0026rsquo; outputs into instructor-led teaching and assessment.\u003c/p\u003e \u003cp\u003eWidespread adoption will also depend on cost and infrastructure. Premium models such as ChatGPT-5 and Gemini 2.5 Flash often require paid subscriptions or institutional licenses, which may limit uptake in lower-resource settings. Computing demands, data privacy regulations, and internet connectivity further constrain global accessibility. Careful attention to pricing structures, data security, and open-source alternatives will be critical to ensure equitable integration.\u003c/p\u003e \u003cp\u003eFinally, high MCQ accuracy does not guarantee clinically valid reasoning. LLMs may reach correct answers through statistical cues rather than genuine understanding. Evaluating confidence calibration, analyzing generated rationales, and comparing reasoning chains with expert logic will be vital for establishing trust.\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eStrengths and Limitations\u003c/h2\u003e \u003cp\u003eKey strengths of this study include a large, diverse set of surgical MCQs and the application of rigorous statistical methods, including McNemar\u0026rsquo;s tests and Cohen\u0026rsquo;s kappa, to quantify model agreement. Stratification by question type and length adds a level of granularity rarely reported in prior LLM evaluations. Nevertheless, several limitations warrant consideration. MCQs reward pattern recognition and may not fully capture deeper clinical reasoning or operative judgment. Image-based questions were excluded, potentially underestimating the advantage of newer multimodal models. Only a single zero-shot prompt was used; alternative strategies might yield different results. The performance of surgical residents or trainees was not included, limiting the educational context. The analysis reflects a single time point and does not account for longitudinal stability or model drift. Accuracy assumes correctness of the provided key.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eImplications and Future Directions\u003c/h2\u003e \u003cp\u003eThe high accuracy of ChatGPT-5 and Gemini 2.5 Flash suggests potential utility as formative assessment tools, question-answer tutors, or aids in MCQ generation for surgical education. However, their consistent deficits in surgical technique raise caution against using these models for procedural instruction without human oversight. Future research should examine hybrid assessment methods that include free-text, simulations, or image-based tasks, optimize prompts, and use chain-of-thought reasoning to improve procedural accuracy, incorporate multimodal inputs like operative images or videos as new models develop visual reasoning abilities, and compare results to human trainees to set meaningful performance benchmarks.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eNewer LLMs, ChatGPT-5 and Gemini 2.5 Flash, significantly outperform ChatGPT-4o on a diverse set of surgical MCQs and achieve substantial agreement with expert answer keys. Their performance is consistent across surgical subspecialties and largely independent of question length, yet all models exhibit persistent deficits in surgical-technique reasoning. These findings support the cautious integration of advanced LLMs into surgical education and assessment while underscoring the need for multimodal approaches, refined prompting, and ongoing human oversight.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eShahab Sheikhalishahi and Alireza Haddadi contributed to data collection. Farzad Rafiei, Shahab Sheikhalishahi, Alireza Haddadi, and Saina Sadeghipour contributed to the conceptualization, methodology, drafting of the original manuscript, and its review and editing. Farzad Rafiei performed the data analysis. Seyed Masoud Hosseini contributed to supervision, as well as the review and editing of the manuscript. All authors read and approved the final version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset of 2025 Iranian surgical medicine subspecialty board questions is publicly available from the official Sanjesh website (https://sanjeshp.ir). Any additional processed data generated and analyzed during the current study are available from the corresponding authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAster A, Laupichler MC, Rockwell-Kollmann T, Masala G, Bala E, Raupach T. ChatGPT and Other Large Language Models in Medical Education \u0026mdash; Scoping Literature Review. Medical Science Educator. 2025;35(1):555-67.\u003c/li\u003e\n\u003cli\u003eBusch F, Hoffmann L, Rueger C, van Dijk EHC, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine. 2025;5(1):26.\u003c/li\u003e\n\u003cli\u003eVrdoljak J, Boban Z, Vilović M, Kumrić M, Božić J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare (Basel). 2025;13(6).\u003c/li\u003e\n\u003cli\u003eLin CR, Chen YJ, Tsai PA, Hsieh WY, Tsai SHL, Fu TS, et al. Multiple large language models versus clinical guidelines for postmenopausal osteoporosis: a comparative study of ChatGPT-3.5, ChatGPT-4.0, ChatGPT-4o, Google Gemini, Google Gemini Advanced, and Microsoft Copilot. Arch Osteoporos. 2025;20(1):120.\u003c/li\u003e\n\u003cli\u003eClark KR. Comparative Analysis of LLMs\u0026apos; Performance On a Practice Radiography Certification Exam. Radiol Technol. 2025;96(5):334-42.\u003c/li\u003e\n\u003cli\u003eKhan AA, Yunus R, Sohail M, Rehman TA, Saeed S, Bu Y, et al. Artificial Intelligence for Anesthesiology Board-Style Examination Questions: Role of Large Language Models. J Cardiothorac Vasc Anesth. 2024;38(5):1251-9.\u003c/li\u003e\n\u003cli\u003ePassby L, Jenko N, Wernham A. Performance of ChatGPT on Specialty Certificate Examination in Dermatology multiple-choice questions. Clin Exp Dermatol. 2024;49(7):722-7.\u003c/li\u003e\n\u003cli\u003eOh N, Choi GS, Lee WY. ChatGPT goes to the operating room: evaluating GPT-4 performance and its potential in surgical education and training in the era of large language models. Ann Surg Treat Res. 2023;104(5):269-73.\u003c/li\u003e\n\u003cli\u003eElzayyat M, Mohammad JN, Zaqout S. Assessing LLM-generated vs. expert-created clinical anatomy MCQs: a student perception-based comparative study in medical education. Med Educ Online. 2025;30(1):2554678.\u003c/li\u003e\n\u003cli\u003eLu Y, Aleta A, Du C, Shi L, Moreno Y. LLMs and generative agent-based models for complex systems research. Phys Life Rev. 2024;51:283-93.\u003c/li\u003e\n\u003cli\u003eChen Y, Huang X, Yang F, Lin H, Lin H, Zheng Z, et al. Performance of ChatGPT and Bard on the medical licensing examinations varies across different cultures: a comparison study. BMC Med Educ. 2024;24(1):1372.\u003c/li\u003e\n\u003cli\u003eYu E, Chu X, Zhang W, Meng X, Yang Y, Ji X, et al. Large Language Models in Medicine: Applications, Challenges, and Future Directions. Int J Med Sci. 2025;22(11):2792-801.\u003c/li\u003e\n\u003cli\u003eZhai C, Wibowo S, Li LD. The effects of over-reliance on AI dialogue systems on students\u0026apos; cognitive abilities: a systematic review. Smart Learning Environments. 2024;11(1):28.\u003c/li\u003e\n\u003cli\u003eSchwartz R, Schwartz R, Vassilev A, Greene K, Perine L, Burt A, et al. Towards a standard for identifying and managing bias in artificial intelligence: US Department of Commerce, National Institute of Standards and Technology \u0026hellip;; 2022.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Surgery, Artificial Intelligence (AI), ChatGPT, Gemini","lastPublishedDoi":"10.21203/rs.3.rs-8080315/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8080315/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis study evaluated the performance of three large language models, including ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash, on 532 Persian multiple-choice questions from the 2025 Iranian surgical subspecialty board examinations. Questions spanned five domains: Pediatric, Cardiovascular, Vascular and Endovascular, Thoracic, and Plastic \u0026amp; Reconstructive Surgery. Using standardized prompts, we assessed overall accuracy, variation across subspecialties and question types, and the effect of question length. ChatGPT-5 and Gemini 2.5 Flash achieved higher accuracy (73.3% and 73.9%) than ChatGPT-4o (68.2%). Agreement with the official key was substantial for Gemini 2.5 Flash (κ\u0026thinsp;=\u0026thinsp;0.651) and ChatGPT-5 (κ\u0026thinsp;=\u0026thinsp;0.642), and moderate to substantial for ChatGPT-4o (κ\u0026thinsp;=\u0026thinsp;0.575). Model performance was stable across subspecialties, but all three showed lower accuracy on surgical technique questions compared with clinical scenarios or basic science items. Question length did not affect ChatGPT-5 or Gemini 2.5 Flash, while longer stems reduced ChatGPT-4o\u0026rsquo;s performance. These findings indicate that newer LLMs provide measurable improvements in surgical question answering, though persistent limitations in procedural reasoning suggest the need for careful integration and further multimodal development.\u003c/p\u003e","manuscriptTitle":"Benchmarking Large Language Models on Persian Surgical Subspecialty Board Examinations: A Comparative Study of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-19 11:35:13","doi":"10.21203/rs.3.rs-8080315/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-06T23:50:52+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-06T10:29:46+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"238356899618282852723789995260061874880","date":"2026-03-02T10:28:48+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"86578774034924550502127756998800134619","date":"2026-03-02T06:16:42+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-23T17:43:54+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"106737375200992459729402047939030280132","date":"2026-02-14T09:43:41+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-13T10:45:53+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-09T10:00:34+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-11-26T19:41:48+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-11-19T18:13:11+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-11-15T08:23:50+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"bcefecf2-9630-4e53-9ccd-5e91ed48c574","owner":[],"postedDate":"February 19th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":62953508,"name":"Health sciences/Health care"},{"id":62953509,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-04-30T10:25:22+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-19 11:35:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8080315","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8080315","identity":"rs-8080315","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.