Performance of Vision–Language Models Compared with 252 Medical Students on Text-only and Image-based Dermatology Examinations

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores > 95 vs. 84.9; p < 0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and ChatGPT-5 achieved higher scores than students, whereas students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. Medical students demonstrated the smallest performance gap between text-only and image-based components, indicating greater cross-modal consistency. Sub-domain analyses revealed that some models achieved accurate visual description and diagnosis but showed reduced performance in etiological and treatment reasoning. Gemini 3 Pro exhibited the highest overall accuracy and the lowest output variability across repeated evaluations. These findings indicate that while VLMs excel in text-based dermatologic assessment, multimodal competence remains uneven and model-dependent, supporting their use as complementary rather than standalone tools in dermatology education.
Full text 123,667 characters · extracted from preprint-html · click to expand
Performance of Vision–Language Models Compared with 252 Medical Students on Text-only and Image-based Dermatology Examinations | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Performance of Vision–Language Models Compared with 252 Medical Students on Text-only and Image-based Dermatology Examinations Ozan Erdem, Abdurrahim Yilmaz, Ahmet Sait Sahin, Bugra Burc Dagtas, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8480126/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores > 95 vs. 84.9; p < 0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and ChatGPT-5 achieved higher scores than students, whereas students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. Medical students demonstrated the smallest performance gap between text-only and image-based components, indicating greater cross-modal consistency. Sub-domain analyses revealed that some models achieved accurate visual description and diagnosis but showed reduced performance in etiological and treatment reasoning. Gemini 3 Pro exhibited the highest overall accuracy and the lowest output variability across repeated evaluations. These findings indicate that while VLMs excel in text-based dermatologic assessment, multimodal competence remains uneven and model-dependent, supporting their use as complementary rather than standalone tools in dermatology education. Biological sciences/Computational biology and bioinformatics Health sciences/Diseases Health sciences/Health care Health sciences/Medical research Vision-Language Models Medical Education Dermatology Medical Students Medical Examinations Large Language Models Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 INTRODUCTION Artificial intelligence (AI) is increasingly influencing diagnostic workflows and educational methodologies in medicine. 1 – 4 This evolution has been driven by rapid advances in large language models (LLMs) and, more recently, vision–language models (VLMs), which are designed to process and synthesize complex biomedical data. 5 – 10 Dermatology, an inherently visual specialty, represents a particularly relevant setting for this multimodal development. Because dermatologic diagnosis relies on integrating visual cues with clinical context to capture fine pathophysiological nuances, the field provides a useful testbed for evaluating whether VLMs can approximate the visual interpretation and integrative reasoning required in clinical practice. 11 – 13 The current literature includes numerous evaluations of the performance of LLMs and VLMs on national medical examinations and specialty board assessments, in which models have been reported to reach or exceed human passing thresholds. 14 – 18 Comparable findings have also been reported in non-English national medical examinations, indicating that the performance of LLMs and VLMs is not confined to English-language assessment settings. 19 – 24 A smaller, but growing, body of work has also directly compared AI performance with that of medical students and residents across different specialties. 25 – 30 However, most benchmarks rely predominantly on text-only questions, leaving the capacity of these models to perform complex visual analysis and to recognize subtle clinical nuances insufficiently examined. 31 This limitation is particularly evident in visually intensive specialties such as dermatology, where diagnostic reasoning depends heavily on the interpretation and integration of image-based information. 32 – 36 Against this background, we conducted a faculty-curated evaluation of contemporary VLMs in undergraduate dermatology education using standardized clerkship examinations at a Turkish medical school. We compared four state-of-the-art models (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on examinations combining text-only multiple-choice, multiple-select, and matching questions (60% of the total score) with image-based, structured open-ended questions (40%) spanning seven domains of dermatologic reasoning and practice. Using expert-validated answer keys, predefined scoring rubrics, and comparative statistical analyses, we aimed to characterize differences in text-only and image-based performance between VLMs and students and to explore domain-specific patterns within a real educational workflow, thereby providing insight into the current strengths and limitations of multimodal AI in dermatology. RESULTS Exam Difficulty Levels and Data Distribution To characterize baseline examination difficulty, student performance across the 10 dermatology clerkship exam sessions was analyzed. Mean student scores in the text-only component ranged from 78.46 ± 9.04 (April 2024) to 91.93 ± 4.66 (February 2024). In the image-based component, mean scores ranged from 72.19 ± 13.70 (December 2024) to 89.41 ± 6.43 (November 2023). Differences in student performance between exam sessions were statistically significant for both text-only and image-based components (Kruskal–Wallis test, p < 0.001 for both). Accordingly, pooled analyses across all 10 examination sessions were performed for subsequent comparisons between VLMs and medical students. Performance on Text-only Examination In the pooled text-only examination, all evaluated VLMs achieved higher mean scores than the medical student cohort (Table 1 ). Students achieved a mean score of 84.91 ± 9.04, whereas all VLMs exceeded 95 points on average (Fig. 1 A). Table 1 Performance of vision–language models and medical students across text-only, image-based, and overall examination modalities. Group Text-only Score Image-based Score Overall Score Gemini 3 Pro 98.45 ± 1.42 93.94 ± 3.41 96.64 ± 1.54 ChatGPT-5 95.19 ± 2.27 87.24 ± 6.97 92.01 ± 2.69 ChatGPT-4o 95.43 ± 1.98 77.36 ± 5.40 88.20 ± 2.01 Gemini 2.5 Flash 96.45 ± 1.42 70.34 ± 5.08 86.00 ± 2.17 Medical Students 84.91 ± 9.04 82.94 ± 9.69 84.12 ± 7.99 Footnote : Data are presented as mean scores ± standard deviation, calculated across ten examination sessions. Overall scores represent a weighted average of text-only (60%) and image-based (40%) components. A Kruskal–Wallis test demonstrated a significant difference among groups (p < 0.001). Post-hoc pairwise comparisons using the Mann–Whitney U test with Bonferroni correction showed that all VLMs (Gemini 3 Pro, Gemini 2.5 Flash, ChatGPT-5, and ChatGPT-4o) scored significantly higher than medical students (p < 0.001 for all comparisons). Among the models, Gemini 3 Pro achieved the highest mean score (98.45 ± 1.42), which was significantly higher than ChatGPT-4o (95.43 ± 1.98, p < 0.001) and ChatGPT-5 (95.19 ± 2.27, p < 0.001). No statistically significant difference was observed between ChatGPT-4o and ChatGPT-5 in the text-only component (p = 1.00). Performance on Image-based Examination In the image-based examination, performance differed significantly among groups (Kruskal–Wallis test, p < 0.001; Table 1 ). Gemini 3 Pro achieved the highest mean score (93.94 ± 3.41), followed by ChatGPT-5 (87.24 ± 6.97). Both models scored higher than the medical student cohort (82.94 ± 9.69), with statistically significant differences (p < 0.001 and p = 0.015, respectively). Medical students achieved higher mean image-based scores than ChatGPT-4o (77.36 ± 5.40) and Gemini 2.5 Flash (70.34 ± 5.08). These differences were statistically significant (p < 0.001 for both comparisons). Gemini 2.5 Flash had the lowest mean score among all groups in the image-based component (Fig. 1 B). Analysis by Text-only Question Type Text-only performance was further analyzed according to question format (Table 2 ). In multiple-choice questions, Gemini 3 Pro achieved the highest accuracy (98.5%), followed by Gemini 2.5 Flash (97.9%), ChatGPT-5 (95.6%), and ChatGPT-4o (95.1%). Medical students achieved a mean accuracy of 82.0%. Differences between students and all VLMs were statistically significant (p < 0.001). Table 2 Accuracy of vision–language models and medical students across different text-only question formats. Group Multiple-choice Multiple-selection Matching Gemini 3 Pro 98.52 ± 1.98 98.31 ± 2.35 98.40 ± 3.18 ChatGPT-5 95.94 ± 3.04 96.41 ± 3.84 88.67 ± 6.77 ChatGPT-4o 96.77 ± 2.52 94.92 ± 4.29 89.60 ± 5.82 Gemini 2.5 Flash 96.26 ± 2.29 97.23 ± 2.63 95.73 ± 4.90 Medical Students 81.99 ± 10.73 89.45 ± 8.54 90.19 ± 10.35 Footnote : Data are presented as mean accuracy percentages ± standard deviation. In multiple-selection questions, Gemini 3 Pro (98.3%) and Gemini 2.5 Flash (97.2%) achieved the highest accuracies, followed by ChatGPT-4o (94.8%) and ChatGPT-5 (94.1%). Medical students achieved a mean accuracy of 85.6%. All VLMs scored significantly higher than students (p < 0.001). In matching questions, Gemini 3 Pro achieved the highest accuracy (98.4%), followed by Gemini 2.5 Flash (95.7%). Medical students achieved a mean accuracy of 90.2%, which was higher than ChatGPT-4o (89.6%) and ChatGPT-5 (88.7%). Differences between students and both ChatGPT models were statistically significant (p < 0.05). Analysis by Image-based Sub-domains Performance in image-based questions was analyzed across seven predefined sub-domains (Table 3 ). Gemini 3 Pro achieved the highest accuracy across all sub-domains, particularly in descriptive reporting (99.4%) and diagnosis (96.7%). Table 3 Accuracy of vision–language models across distinct image-based dermatological sub-domains. Sub-domain Gemini 3 Pro ChatGPT-5 ChatGPT-4o Gemini 2.5 Flash Elementary lesion recognition 89.40 ± 7.47 79.90 ± 12.92 79.60 ± 16.44 75.00 ± 15.75 Descriptive reporting 99.40 ± 3.14 99.00 ± 3.64 96.60 ± 8.48 77.60 ± 18.58 Diagnosis 96.67 ± 10.1 96.67 ± 8.65 90.00 ± 14.52 86.40 ± 16.38 Etiology 95.00 ± 15.1 94.80 ± 15.15 61.80 ± 28.76 53.00 ± 35.93 Differential diagnosis 87.80 ± 13.2 73.40 ± 23.33 62.67 ± 26.98 54.73 ± 24.93 Treatment 88.60 ± 7.56 86.20 ± 12.60 79.60 ± 10.68 58.60 ± 23.04 General knowledge 97.92 ± 4.14 86.24 ± 16.92 72.56 ± 11.14 72.48 ± 13.89 Data are presented as mean accuracy percentages ± standard deviation. In the etiology sub-domain, Gemini 3 Pro (95.0%) and ChatGPT-5 (94.8%) achieved higher accuracies than ChatGPT-4o (61.8%) and Gemini 2.5 Flash (53.0%). Differences between model performances in this sub-domain were statistically significant (p < 0.001). In treatment-related questions, Gemini 3 Pro and ChatGPT-5 again achieved higher accuracies compared with ChatGPT-4o, Gemini 2.5 Flash, and the student cohort. Performance patterns across remaining sub-domains are summarized in Table 3 and Fig. 2 B. Comparison of Text-only and Image-based Performance Within-group comparisons between text-only and image-based scores were performed using the Wilcoxon signed-rank test (Fig. 3 ). All groups demonstrated statistically significant differences between modalities (p < 0.05). The largest mean score differences between text-only and image-based components were observed for Gemini 2.5 Flash (26.1 points, p < 0.001) and ChatGPT-4o (18.0 points, p < 0.001). ChatGPT-5 demonstrated a smaller difference of 8.0 points (p < 0.001), while Gemini 3 Pro showed the smallest difference among models (4.5 points, p < 0.001). Notably, medical students demonstrated the most balanced profile with a mean difference of 1.97 points between modalities (p = 0.002). Correlation with Examination Difficulty To assess the relationship between examination difficulty and group performance, mean scores of each VLM were correlated with mean student scores across the 10 exam sessions (Fig. 4 ). In the text-only component, no significant correlation was observed between student scores and VLM scores. In contrast, in the image-based component, positive correlations were observed between student performance and the performance of ChatGPT-4o and Gemini 2.5 Flash, whereas Gemini 3 Pro maintained consistently high performance across sessions. Model Output Variability Model output variability was assessed by analyzing score distributions across five repeated runs for each exam session (Fig. 5 ). Gemini 3 Pro demonstrated the lowest mean standard deviation across repeated runs in both text-only (SD = 0.72) and image-based (SD = 1.70) components, while higher variability was observed for Gemini 2.5 Flash, particularly in the image-based component (SD = 3.23). ChatGPT-4o and ChatGPT-5 demonstrated intermediate levels of variability. DISCUSSION This study provides a structured comparison of contemporary VLMs and medical students within an authentic undergraduate dermatology assessment framework, revealing clear modality-dependent differences in performance. By analyzing both text-only and image-based examination components across multiple exam sessions, we found that strong performance in text-based dermatologic assessments does not consistently translate to image-based reasoning. While several VLMs achieved near-ceiling scores on text-only tasks, performance diverged substantially on image-based questions, with marked variability across model architectures. These findings indicate that multimodal competence in dermatology remains heterogeneous and task-dependent, underscoring the importance of evaluating visual and textual reasoning separately rather than inferring overall capability from text-based performance alone. In the text-only domain, all evaluated VLMs achieved near-ceiling performance, exceeding the mean scores of the medical student cohort across exam sessions. Moreover, model performance remained largely invariant to fluctuations in exam difficulty that significantly affected student scores. These findings are consistent with prior studies demonstrating that LLMs can reach or exceed passing thresholds on national licensing and specialty board examinations, reflecting strong capabilities in factual recall and structured medical knowledge application. 14 , 37 , 38 However, aggregate text-based scores obscure meaningful differences between question formats. While VLMs consistently outperformed students on multiple-choice and multiple-select items, students achieved comparable or slightly higher accuracy than ChatGPT-4o and ChatGPT-5 on matching-type questions, suggesting that certain associative reasoning tasks without contextual scaffolding may still favor human cognitive strategies. The image-based examination component revealed a markedly different performance hierarchy. While Gemini 3 Pro and ChatGPT-5 achieved higher mean scores than medical students, students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. This reversal underscores that visual diagnostic competence remains a challenging domain for some VLMs, despite their strong performance on text-based tasks. Similar modality-dependent performance patterns have been reported in recent evaluations, where accuracy declined for items incorporating images or complex visual contexts. 23 , 31 , 39 Notably, the student cohort exhibited the smallest performance difference between text-only and image-based tasks, indicating a relatively consistent integration of theoretical knowledge and visual interpretation. In contrast, earlier or lighter model variants showed substantial performance declines when transitioning from text to image-based tasks, reflecting limited multimodal consistency. Consequently, text-based accuracy alone is an incomplete proxy for multimodal competence and should be interpreted cautiously when evaluating the readiness of VLMs for visually intensive clinical domains such as dermatology. 40 Sub-domain analyses further clarify the nature of this discrepancy. While most models performed well in tasks emphasizing visual description and pattern recognition, notable limitations emerged in domains requiring etiological reasoning from visual input. For example, ChatGPT-4o and Gemini 2.5 Flash demonstrated reduced accuracy in etiology-related questions despite achieving moderate diagnostic accuracy. This pattern suggests that correct visual labeling does not necessarily coincide with coherent pathophysiological reasoning, a distinction that has also been emphasized in prior analyses of AI-assisted dermatologic diagnosis. 41 In contrast, Gemini 3 Pro and ChatGPT-5 maintained higher performance across etiological and treatment-related domains, indicating improved integration of visual features with underlying clinical knowledge. Model reliability represents an additional and clinically relevant dimension. Repeated exam administrations revealed substantial variability in output stability across models. Gemini 3 Pro exhibited the lowest performance variance across both modalities, whereas Gemini 2.5 Flash showed higher stochastic variability, particularly in image-based tasks. Importantly, such variability may limit the interpretability and pedagogical reliability of model outputs in educational settings. These findings highlight that average accuracy alone is insufficient for evaluating AI systems intended for educational or clinical support, as reproducibility and consistency are essential prerequisites for trust and safe deployment. 42 Several methodological strengths enhance the relevance of this study. The use of original, non-English examination materials provides insight into multilingual model performance in real educational settings. The inclusion of a large human cohort and multiple repeated AI evaluations allows robust statistical comparison and characterization of performance variability. Nevertheless, important limitations should be acknowledged. This was a single-center study reflecting one institutional curriculum, and the examinations, while clinically oriented, remain structured academic assessments. Static images cannot fully capture the complexity of real-world dermatologic evaluation, which often incorporates dermoscopy, palpation, and longitudinal clinical context. Taken together, these findings indicate that while contemporary VLMs achieve very high performance on text-based dermatologic assessments, their image-based diagnostic capabilities remain heterogeneous and strongly model-dependent. Medical students, although achieving lower absolute scores in text-only tasks, display greater consistency across modalities, reflecting the integrative nature of clinical training. Continued improvements in multimodal reasoning are evident in newer model generations; however, careful validation and human oversight remain essential. Rather than replacing human expertise, VLMs may best be positioned as complementary tools in dermatology education, supporting learning and assessment within a controlled, human-in-the-loop framework. METHODS Study Design and Setting This comparative cross-sectional study assessed the performance of VLMs and fifth-year medical students on standardized dermatology clerkship examinations at Istanbul Medeniyet University Faculty of Medicine. The dataset comprised ten consecutive clerkship rotations conducted between September 2023 and January 2025. Both the text-only and image-based examinations were mandatory components of the official clerkship assessment. All analyses were performed using the original Turkish questions and responses to evaluate performance in the language of instruction. Participants A total of 252 fifth-year medical students participated across ten clerkship groups (21–29 students per group; mean ≈ 25). Because the examinations were a required part of the rotation, no exclusion criteria were applied. Exam Structure and Scoring Across the ten exam sessions, the dataset included 500 text-only items and 200 image-based items. Each session consisted of: Text-only examination (50 items): 31 single-answer multiple-choice (Q1–31), 13 multiple-select (Q32–44), and 6 matching questions (Q45–50). Image-based examination (20 items): structured open-ended questions spanning seven sub-domains: elementary lesion recognition, descriptive reporting, diagnosis, etiology, differential diagnosis, treatment, and general dermatological knowledge. Text-only items were graded using predefined answer keys, whereas image-based items were graded using weighted rubrics. The overall examination score was calculated as a weighted combination of modalities (text-only 60%, image-based 40%). The detailed exam structure and scoring framework is summarized in Table 4 . Table 4 Structure and scoring scheme of the dermatology clerkship examinations. Exam component Subcategory (question range) No. of items Scoring scheme Max score Text-only exam Multiple-choice (Q1–31) 31 2 points per correct answer 62 Multiple-select (Q32–44) 13 (Correct selections ÷ 3) × 2; max 2 points per item; invalid (0 points) if > 3 choices selected 26 Matching (Q45–50) 6 (Correct matches ÷ 5) × 2; max 2 points per item 12 Total 50 — 100 Image-based exam Elementary lesion recognition (Q1–4) 4 Weighted rubric; max 5 points per item 20 Descriptive reporting (Q5–6) 2 Weighted rubric; max 5 points per item 10 Diagnosis (Q7–9) 3 Single-answer; max 5 points per item 15 Etiology (Q10–11) 2 Single-answer; max 5 points per item 10 Differential diagnosis (Q12–13) 2 Weighted rubric; max 5 points per item 10 Treatment (Q14–15) 2 Weighted rubric or single-answer; max 5 points per item 10 General dermatological knowledge (Q16–20) 5 Weighted rubric or single-answer; max 5 points per item 25 Total 20 — 100 Footnote : Text-only and image-based components were each scaled to a maximum of 100 points. For multiple-select items, partial credit was calculated as (number of correct selections ÷ 3) × 2 (maximum 2 points per item); responses selecting more than three options were scored as 0. For matching items, partial credit was calculated as (number of correct matches ÷ 5) × 2 (maximum 2 points per item). ‘Weighted rubric’ indicates predefined point allocations (up to 5 points per item) based on faculty-approved grading criteria (e.g., awarding credit for specific expected elements within an open-ended response). Overall score was computed as a weighted average: 0.60 × text-only score + 0.40 × image-based score. AI Models and Testing Procedure Four VLMs developed by OpenAI (ChatGPT-4o, ChatGPT-5) and Google (Gemini 2.5 Flash, Gemini 3 Pro) were evaluated between May and October 2025. Model testing used standardized, format-specific Turkish prompts applied identically across all models and examination items. Text-only items were submitted in grouped batches with explicit formatting instructions (multiple-choice, multiple-select, or matching). Image-based items were submitted individually and required structured open-ended responses with predefined constraints on the type and number of expected answers. A representative testing workflow is shown in Fig. 6 . To account for stochastic variability in generative outputs, each exam session was repeated five times per model. For transparency, standardized prompt templates and representative examples are provided in Supplementary Information I, and the full content of the first evaluated examination (22 September 2023) with English translations is provided in Supplementary Information II–III, together with the corresponding expert-validated answer keys. Gold-standard Answers and Evaluation of Model Outputs Faculty members from relevant subspecialty areas authored the examination items. All questions were compiled, reviewed, and approved by the clerkship coordinator. Gold-standard answers (text-only) and grading rubrics (image-based) were prepared by the item authors and verified by the clerkship coordinator. VLM responses were scored by the clerkship coordinator and a senior dermatology resident through collaborative review with task sharing, reflecting the routine grading workflow used in the clerkship. No independent parallel scoring, averaging, or formal consensus procedure was applied. Availability of Student Image-based Sub-domain Data Student image-based data were obtained retrospectively from official clerkship records. While total image-based examination scores were available for all students, sub-domain–level rubric components were not recorded in the examination system at the time of assessment. Accordingly, sub-domain analyses for image-based performance could be conducted only for VLMs, for which rubric-based scoring was prospectively recorded at the item level. Statistical Analysis All statistical analyses were performed in Python (version 3.10) using the pandas , numpy , and scipy libraries. Data visualization was conducted using matplotlib and seaborn to generate violin plots, radar charts, and slope graphs. Normality of score distributions was assessed using the Shapiro–Wilk test. As several variables deviated from normality, non-parametric statistical methods were applied. Overall between-group differences were evaluated using the Kruskal–Wallis test, followed by pairwise Mann–Whitney U tests for relevant group comparisons (students versus each model and between-model comparisons), with Bonferroni correction applied for multiple testing. To assess within-group differences between text-only and image-based performance (the modality gap), the Wilcoxon signed-rank test was used. Correlations between mean student scores and model scores across examination sessions were examined using Spearman’s rank correlation coefficient. Model output variability was quantified descriptively using the standard deviation of total scores across five repeated runs per model and examination session. Statistical significance was defined as a two-sided p value < 0.05. Declarations Funding information: This article has no funding source. Conflicts of Interest: The authors have no conflict of interest to declare. Ethical Approval: Reviewed and approved by the Ethics Committee of Prof. Dr. Suleyman Yalcin City Hospital (Approval No.: 2025/0263) and Deanship of the Faculty of Medicine, Istanbul Medeniyet University (Approval No.: E-28298836-100-2500069724). This study was conducted retrospectively, and the requirement for informed consent for participants was waived by the Ethics Committee and confirmed by the Faculty of Medicine Deanship. Ethics statement: The study was conducted in accordance with the principles of the Declaration of Helsinki. Data availability statement : The data that support the findings of this study are available from the corresponding author upon reasonable request. Author contribution: O.E. and A.Y. conceptualized and designed the study and developed the methodology. Formal analysis was performed by O.E., A.S.Ş., and A.Y. Investigation was carried out by O.E. and A.S.Ş. Resources were provided by O.E., V.A.E., M.A.K., and M.S.G. O.E. prepared the visualizations. The original draft of the manuscript was written by O.E. and A.Y. Manuscript review and editing were performed by E.G. and B.B.D. M.S.G. supervised the study. All authors read and approved the final manuscript. Acknowledgement: Abdurrahim Yilmaz has been funded by the President’s PhD Scholarships at Imperial College London, which solely supported his academic studies. References Esteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature 542 , 115–118 (2017). Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25 , 44–56 (2019). Gordon, M. et al. A scoping review of artificial intelligence in medical education: BEME Guide 84. Med. Teach. 46 , 446–470 (2024). Vrdoljak, J., Boban, Z., Vilović, M., Kumrić, M. & Božić, J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare 13 , 603 (2025). Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29 , 1930–1940 (2023). Singhal, K. et al. Large language models encode clinical knowledge. Nature 620 , 172–180 (2023). Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616 , 259–265 (2023). Hartsock, I. & Rasool, G. Vision-language models for medical report generation and visual question answering: a review. Front. Artif. Intell. 7 , 1430984 (2024). Krones, F., Marikkar, U., Parsons, G., Szmul, A. & Mahdi, A. Review of multimodal machine learning approaches in healthcare. Inf. Fusion . 114 , 102690 (2025). Wang, Z. et al. A perspective for adapting generalist AI to specialized medical AI applications and their challenges. Npj Digit. Med. 8 , 429 (2025). Yan, S. et al. A multimodal vision foundation model for clinical dermatology. Nat. Med. 31 , 2691–2702 (2025). Zhou, J. et al. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nat. Commun. 15 , 5649 (2024). Yilmaz, A. et al. Resource-efficient medical vision language model for dermatology via a synthetic data generation framework. 05.17.25327785 Preprint at (2025). https://doi.org/10.1101/2025.05.17.25327785 (2025). Kung, T. H. et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit. Health . 2 , e0000198 (2023). Gilson, A. et al. How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment. JMIR Med. Educ. 9 , e45312 (2023). Brin, D. et al. Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments. Sci. Rep. 13 , 16492 (2023). Sadeq, M. A. et al. AI chatbots show promise but limitations on UK medical exam questions: a comparative performance study. Sci. Rep. 14 , 18859 (2024). Casals-Farre, O. et al. Assessing ChatGPT 4.0’s Capabilities in the United Kingdom Medical Licensing Examination (UKMLA): A Robust Categorical Analysis. Sci. Rep. 15 , 13031 (2025). Kim, H. J. et al. Performance evaluation of large language models on Korean medical licensing examination: a three-year comparative analysis. Sci. Rep. 15 , 36082 (2025). Jung, L. B. et al. ChatGPT Passes German State Examination in Medicine With Picture Questions Omitted. Dtsch. Arzteblatt Int. 120 , 373–374 (2023). Takagi, S., Watari, T., Erabi, A. & Sakaguchi, K. Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: Comparison Study. JMIR Med. Educ. 9 , e48002 (2023). Rojas, M., Rojas, M., Burgess, V., Toro-Pérez, J. & Salehi, S. Exploring the Performance of ChatGPT Versions 3.5, 4, and 4 With Vision in the Chilean Medical Licensing Examination: Observational Study. JMIR Med. Educ. 10 , e55048–e55048 (2024). Miyazaki, Y. et al. Performance of ChatGPT-4o on the Japanese Medical Licensing Examination: Evalution of Accuracy in Text-Only and Image-Based Questions. JMIR Med. Educ. 10 , e63129–e63129 (2024). Luo, D. et al. Evaluating the performance of GPT-3.5, GPT-4, and GPT-4o in the Chinese National Medical Licensing Examination. Sci. Rep. 15 , 14119 (2025). Guerra, G. A. et al. GPT-4 Artificial Intelligence Model Outperforms ChatGPT, Medical Students, and Neurosurgery Residents on Neurosurgery Written Board-Like Questions. World Neurosurg. 179 , e160–e165 (2023). Massey, P. A., Montgomery, C. & Zhang, A. S. Comparison of ChatGPT–3.5, ChatGPT-4, and Orthopaedic Resident Performance on Orthopaedic Assessment Examinations. J. Am. Acad. Orthop. Surg. 31 , 1173–1179 (2023). Roos, J., Kasapovic, A., Jansen, T. & Kaczmarczyk, R. Artificial Intelligence in Medical Education: Comparative Analysis of ChatGPT, Bing, and Medical Students in Germany. JMIR Med. Educ. 9 , e46482 (2023). Meyer, A., Riese, J. & Streichert, T. Comparison of the Performance of GPT-3.5 and GPT-4 With That of Medical Students on the Written German Medical Licensing Examination: Observational Study. JMIR Med. Educ. 10 , e50965 (2024). Bahir, D. et al. Gemini AI vs. ChatGPT: A comprehensive examination alongside ophthalmology residents in medical knowledge. Graefes Arch. Clin. Exp. Ophthalmol. https://doi.org/10.1007/s00417-024-06625-4 (2024). Zengin, A., Ulfanov, O., Bag, Y. M. & Ulas, M. Artificial Intelligence Versus Medical Students in General Surgery Exam. Indian J. Surg. 87 , 68–73 (2025). Yang, X. & Chen, W. The performance of ChatGPT on medical image-based assessments and implications for medical education. BMC Med. Educ. 25 , 1192 (2025). Behrmann, J. et al. Chat generative pre-trained transformer’s performance on dermatology-specific questions and its implications in medical education. J. Med. Artif. Intell. 6 , 16–16 (2023). Passby, L., Jenko, N. & Wernham, A. Performance of ChatGPT on Specialty Certificate Examination in Dermatology multiple-choice questions. Clin. Exp. Dermatol. 49 , 722–727 (2024). Fan, K. S. & Fan, K. H. Dermatological Knowledge and Image Analysis Performance of Large Language Models Based on Specialty Certificate Examination in Dermatology. Dermato 4 , 124–135 (2024). Göçer Gürok, N. & Öztürk, S. The Performance of AI in Dermatology Exams: The Exam Success and Limits of ChatGPT. J. Cosmet. Dermatol. 24 , e70244 (2025). Atılan, A. U. & Çetin, N. Benchmarking Large Language Models on the Turkish Dermatology Board Exam: A Comparative Multilingual Analysis. Turk. J. Dermatol. 19 , 126–133 (2025). Bicknell, B. T. et al. ChatGPT-4 Omni Performance in USMLE Disciplines and Clinical Skills: Comparative Analysis. JMIR Med. Educ. 10 , e63430 (2024). Brin, D. et al. How GPT models perform on the United States medical licensing examination: a systematic review. Discov Appl. Sci. 6 , 500 (2024). Liu, M. et al. Evaluating the Effectiveness of advanced large language models in medical Knowledge: A Comparative study using Japanese national medical examination. Int. J. Med. Inf. 193 , 105673 (2025). Jin, Q. et al. Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine. Npj Digit. Med. 7 , 190 (2024). Yu, Z., Xin, C., Yu, Y., Xia, J. & Han, L. AI dermatology: Reviewing the frontiers of skin cancer detection technologies. Intell. Oncol. 1 , 89–104 (2025). Han, H. Challenges of reproducible AI in biomedical data science. BMC Med. Genomics 18 , 8, (2025). s12920-024-02072–6. Additional Declarations No competing interests reported. Supplementary Files SupplementaryInformation1.pdf SupplementaryInformation2.pdf SupplementaryInformation3.pdf Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 30 Apr, 2026 Reviewers agreed at journal 27 Apr, 2026 Reviewers agreed at journal 24 Apr, 2026 Reviews received at journal 22 Apr, 2026 Reviewers agreed at journal 13 Apr, 2026 Reviewers invited by journal 05 Apr, 2026 Editor assigned by journal 31 Mar, 2026 Editor invited by journal 05 Jan, 2026 Submission checks completed at journal 02 Jan, 2026 First submitted to journal 02 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8480126","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":618713967,"identity":"2db42572-6f2a-48bd-b547-949a54078d16","order_by":0,"name":"Ozan Erdem","email":"","orcid":"","institution":"Istanbul Medeniyet University","correspondingAuthor":false,"prefix":"","firstName":"Ozan","middleName":"","lastName":"Erdem","suffix":""},{"id":618713968,"identity":"df2d5b4c-f60a-4643-bb8b-b10256783b1b","order_by":1,"name":"Abdurrahim Yilmaz","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABE0lEQVRIie3RMUvDQBTA8RcOrg6nWU9Q+xVecRJK81V6BJxSdLfgQaCT4loR/AyCUDq+epAuFWe3BFcHSxeXUi9tQEEvXR3uPxy8wI/L4wB8vn8a5fYIw5SonDgwvf4s60jXHvvDrFuRYDuBkiAlWI1bSKhFi1TSiYBm86fXcftsrzHR8NkHdav/JpIEkhrFLEivH0xvdnoyEEoHVxmoO+cia8I4g2dLBgY5KA27GtS9QzQ35FJwSHJLVsjDQgfLGoIbYqSABCwh5FJpVt7i+rGW4eeWTFHKDO0usSWFNgeZPHatfzRNHz/mo4soeknfFr1xB5s38aR477cPh+S4hu3g91C9CFD9Qzby38Tn8/l8P/sCBYNhi5EZynIAAAAASUVORK5CYII=","orcid":"","institution":"Imperial College London","correspondingAuthor":true,"prefix":"","firstName":"Abdurrahim","middleName":"","lastName":"Yilmaz","suffix":""},{"id":618713969,"identity":"4c5813c6-0a90-4dcd-9c80-553ae69842fb","order_by":2,"name":"Ahmet Sait Sahin","email":"","orcid":"","institution":"Istanbul Medeniyet University","correspondingAuthor":false,"prefix":"","firstName":"Ahmet","middleName":"Sait","lastName":"Sahin","suffix":""},{"id":618713970,"identity":"a601c047-e548-49c1-9ddf-f7fcc4369226","order_by":3,"name":"Bugra Burc Dagtas","email":"","orcid":"","institution":"İstanbul Eğitim ve Araştırma Hastanesi","correspondingAuthor":false,"prefix":"","firstName":"Bugra","middleName":"Burc","lastName":"Dagtas","suffix":""},{"id":618713971,"identity":"ec55d8e1-d6eb-45c7-9df8-cddf895d8b2c","order_by":4,"name":"Ece Gokyayla","email":"","orcid":"","institution":"Ege University","correspondingAuthor":false,"prefix":"","firstName":"Ece","middleName":"","lastName":"Gokyayla","suffix":""},{"id":618713972,"identity":"3d88192c-577b-481d-aa32-f62e8cd8d0e3","order_by":5,"name":"Melek Aslan Kayıran","email":"","orcid":"","institution":"Istanbul Medeniyet University","correspondingAuthor":false,"prefix":"","firstName":"Melek","middleName":"Aslan","lastName":"Kayıran","suffix":""},{"id":618713973,"identity":"36b58142-c70c-498d-a833-a391cb3f7324","order_by":6,"name":"Vefa Aslı Erdemir","email":"","orcid":"","institution":"Istanbul Medeniyet University","correspondingAuthor":false,"prefix":"","firstName":"Vefa","middleName":"Aslı","lastName":"Erdemir","suffix":""},{"id":618713974,"identity":"06a1b39e-4fc9-4927-a633-90052d85a2c4","order_by":7,"name":"Mehmet Salih Gurel","email":"","orcid":"","institution":"Istanbul Medeniyet University","correspondingAuthor":false,"prefix":"","firstName":"Mehmet","middleName":"Salih","lastName":"Gurel","suffix":""}],"badges":[],"createdAt":"2025-12-30 10:23:06","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8480126/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8480126/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106724965,"identity":"d09484e3-9f90-4313-8fa9-f00dbaf961c0","added_by":"auto","created_at":"2026-04-12 18:30:48","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":132061,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of performance scores across text-only, image-based, and overall assessment modalities. Violin plots display score distributions for each participant group. The width of each violin represents kernel density estimation. Internal box plots indicate the median (central line), interquartile range (box limits), and whiskers (1.5× interquartile range). White diamond markers denote mean scores. \u003cstrong\u003eLeft panel:\u003c/strong\u003e text-only examination scores. \u003cstrong\u003eMiddle panel:\u003c/strong\u003eimage-based examination scores. \u003cstrong\u003eRight panel:\u003c/strong\u003e overall examination scores combining text-only and image-based components.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/6105998406d9258fad47cffb.png"},{"id":106725050,"identity":"a055cff4-fd10-46d7-82e3-d949d1d81fb8","added_by":"auto","created_at":"2026-04-12 18:31:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":243288,"visible":true,"origin":"","legend":"\u003cp\u003eComparative performance of vision–language models and medical students across text-only and image-based assessment domains. \u003cstrong\u003e(A)\u003c/strong\u003e Mean accuracy percentages for text-only examinations stratified by question type (multiple choice, multiple selection, and matching). \u003cstrong\u003e(B)\u003c/strong\u003e Radar chart showing mean accuracy across seven image-based sub-domains: elementary lesion recognition, descriptive reporting, diagnosis, etiology, differential diagnosis, treatment, and general dermatological knowledge. Sub-domain–level image-based performance is shown for vision–language models only, as item-level rubric data were not available for medical students.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/2f08384b4016508da3cb995b.png"},{"id":106727092,"identity":"c0ba5748-4ce0-4716-a0f0-b81a69ab51ff","added_by":"auto","created_at":"2026-04-12 18:38:06","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":153511,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of text-only and image-based performance within each participant group. Lines connect mean text-only scores (left) to mean image-based scores (right). The vertical distance between points represents the magnitude of score difference between modalities for each group.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/b9a6948c11f208548a42c6bb.png"},{"id":106545315,"identity":"6bd8063d-8ed6-4f61-a826-77e6ed83ec2d","added_by":"auto","created_at":"2026-04-09 16:45:19","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":156484,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation between examination difficulty and group performance across 10 exam sessions. Scatter plots display the relationship between mean medical student scores for each exam session (x-axis) and corresponding mean scores of vision–language models (y-axis). \u003cstrong\u003e(A)\u003c/strong\u003e Text-only examination scores. \u003cstrong\u003e(B)\u003c/strong\u003eImage-based examination scores. Trend lines represent linear regression for each model.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/6be7f96b6c10c9d736f6def8.png"},{"id":106725030,"identity":"384d98f6-e360-47ca-b515-d89cc19424b5","added_by":"auto","created_at":"2026-04-12 18:31:05","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":87857,"visible":true,"origin":"","legend":"\u003cp\u003eModel output variability across repeated evaluations. Bar chart displays the mean standard deviation of scores across five repeated runs for each vision–language model in text-only and image-based examinations. Lower standard deviation values indicate lower score variability across runs.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/de2c1473ed1ffb0ab2ba918f.png"},{"id":106724950,"identity":"038df346-e772-44c5-8022-d590ed323303","added_by":"auto","created_at":"2026-04-12 18:30:40","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":633041,"visible":true,"origin":"","legend":"\u003cp\u003eRepresentative examples from the text-only \u003cstrong\u003e(left)\u003c/strong\u003e and image-based \u003cstrong\u003e(right) \u003c/strong\u003ecomponents of the dermatology clerkship examinations, illustrating the standardized prompt templates and example model outputs. Prompt templates were format-specific and instructed models to return answers in a constrained output format (e.g., letter-only responses for text-only items and short, structured responses for image-based items). \u003cem\u003eAll analyses were performed using the original Turkish questions and responses; English translations are provided for illustrative purposes only.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/8de89a2411d2c61fc54ae0f0.png"},{"id":106728147,"identity":"d23fec4c-d245-41cd-bb00-d474da63b6f4","added_by":"auto","created_at":"2026-04-12 18:41:59","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2534088,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/56fd5d58-fca9-492b-bc2f-1c44ea747a55.pdf"},{"id":106724970,"identity":"a2aeb352-259e-4878-91b7-975ff58a0e90","added_by":"auto","created_at":"2026-04-12 18:30:48","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":136954,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/b5e3e2584e16d43556fc150a.pdf"},{"id":106545310,"identity":"38c33a6b-e098-4163-b4eb-75fb3e2d1aab","added_by":"auto","created_at":"2026-04-09 16:45:19","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":261213,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/b4abab1a2a1026fc69ce4938.pdf"},{"id":106545314,"identity":"3d4e9086-6eb7-453a-8adf-0573ba866e4c","added_by":"auto","created_at":"2026-04-09 16:45:19","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1610585,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8480126/v1/f532e769dbe9336fc816f6ad.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003ePerformance of Vision–Language Models Compared with 252 Medical Students on Text-only and Image-based Dermatology Examinations\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eArtificial intelligence (AI) is increasingly influencing diagnostic workflows and educational methodologies in medicine.\u003csup\u003e\u003cspan additionalcitationids=\"CR2 CR3\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e This evolution has been driven by rapid advances in large language models (LLMs) and, more recently, vision\u0026ndash;language models (VLMs), which are designed to process and synthesize complex biomedical data.\u003csup\u003e\u003cspan additionalcitationids=\"CR6 CR7 CR8 CR9\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e Dermatology, an inherently visual specialty, represents a particularly relevant setting for this multimodal development. Because dermatologic diagnosis relies on integrating visual cues with clinical context to capture fine pathophysiological nuances, the field provides a useful testbed for evaluating whether VLMs can approximate the visual interpretation and integrative reasoning required in clinical practice.\u003csup\u003e\u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThe current literature includes numerous evaluations of the performance of LLMs and VLMs on national medical examinations and specialty board assessments, in which models have been reported to reach or exceed human passing thresholds.\u003csup\u003e\u003cspan additionalcitationids=\"CR15 CR16 CR17\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e Comparable findings have also been reported in non-English national medical examinations, indicating that the performance of LLMs and VLMs is not confined to English-language assessment settings.\u003csup\u003e\u003cspan additionalcitationids=\"CR20 CR21 CR22 CR23\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e A smaller, but growing, body of work has also directly compared AI performance with that of medical students and residents across different specialties.\u003csup\u003e\u003cspan additionalcitationids=\"CR26 CR27 CR28 CR29\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e However, most benchmarks rely predominantly on text-only questions, leaving the capacity of these models to perform complex visual analysis and to recognize subtle clinical nuances insufficiently examined.\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e This limitation is particularly evident in visually intensive specialties such as dermatology, where diagnostic reasoning depends heavily on the interpretation and integration of image-based information.\u003csup\u003e\u003cspan additionalcitationids=\"CR33 CR34 CR35\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eAgainst this background, we conducted a faculty-curated evaluation of contemporary VLMs in undergraduate dermatology education using standardized clerkship examinations at a Turkish medical school. We compared four state-of-the-art models (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on examinations combining text-only multiple-choice, multiple-select, and matching questions (60% of the total score) with image-based, structured open-ended questions (40%) spanning seven domains of dermatologic reasoning and practice. Using expert-validated answer keys, predefined scoring rubrics, and comparative statistical analyses, we aimed to characterize differences in text-only and image-based performance between VLMs and students and to explore domain-specific patterns within a real educational workflow, thereby providing insight into the current strengths and limitations of multimodal AI in dermatology.\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eExam Difficulty Levels and Data Distribution\u003c/h2\u003e \u003cp\u003eTo characterize baseline examination difficulty, student performance across the 10 dermatology clerkship exam sessions was analyzed. Mean student scores in the text-only component ranged from 78.46\u0026thinsp;\u0026plusmn;\u0026thinsp;9.04 (April 2024) to 91.93\u0026thinsp;\u0026plusmn;\u0026thinsp;4.66 (February 2024). In the image-based component, mean scores ranged from 72.19\u0026thinsp;\u0026plusmn;\u0026thinsp;13.70 (December 2024) to 89.41\u0026thinsp;\u0026plusmn;\u0026thinsp;6.43 (November 2023).\u003c/p\u003e \u003cp\u003eDifferences in student performance between exam sessions were statistically significant for both text-only and image-based components (Kruskal\u0026ndash;Wallis test, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for both). Accordingly, pooled analyses across all 10 examination sessions were performed for subsequent comparisons between VLMs and medical students.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePerformance on Text-only Examination\u003c/h3\u003e\n\u003cp\u003eIn the pooled text-only examination, all evaluated VLMs achieved higher mean scores than the medical student cohort (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Students achieved a mean score of 84.91\u0026thinsp;\u0026plusmn;\u0026thinsp;9.04, whereas all VLMs exceeded 95 points on average (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of vision\u0026ndash;language models and medical students across text-only, image-based, and overall examination modalities.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGroup\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eText-only Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eImage-based Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOverall Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGemini 3 Pro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e98.45\u0026thinsp;\u0026plusmn;\u0026thinsp;1.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e93.94\u0026thinsp;\u0026plusmn;\u0026thinsp;3.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e96.64\u0026thinsp;\u0026plusmn;\u0026thinsp;1.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChatGPT-5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e95.19\u0026thinsp;\u0026plusmn;\u0026thinsp;2.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e87.24\u0026thinsp;\u0026plusmn;\u0026thinsp;6.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e92.01\u0026thinsp;\u0026plusmn;\u0026thinsp;2.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChatGPT-4o\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e95.43\u0026thinsp;\u0026plusmn;\u0026thinsp;1.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e77.36\u0026thinsp;\u0026plusmn;\u0026thinsp;5.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e88.20\u0026thinsp;\u0026plusmn;\u0026thinsp;2.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGemini 2.5 Flash\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e96.45\u0026thinsp;\u0026plusmn;\u0026thinsp;1.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e70.34\u0026thinsp;\u0026plusmn;\u0026thinsp;5.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e86.00\u0026thinsp;\u0026plusmn;\u0026thinsp;2.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMedical Students\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e84.91\u0026thinsp;\u0026plusmn;\u0026thinsp;9.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e82.94\u0026thinsp;\u0026plusmn;\u0026thinsp;9.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e84.12\u0026thinsp;\u0026plusmn;\u0026thinsp;7.99\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003cb\u003eFootnote\u003c/b\u003e: Data are presented as mean scores\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation, calculated across ten examination sessions. Overall scores represent a weighted average of text-only (60%) and image-based (40%) components.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eA Kruskal\u0026ndash;Wallis test demonstrated a significant difference among groups (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Post-hoc pairwise comparisons using the Mann\u0026ndash;Whitney U test with Bonferroni correction showed that all VLMs (Gemini 3 Pro, Gemini 2.5 Flash, ChatGPT-5, and ChatGPT-4o) scored significantly higher than medical students (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for all comparisons).\u003c/p\u003e \u003cp\u003eAmong the models, Gemini 3 Pro achieved the highest mean score (98.45\u0026thinsp;\u0026plusmn;\u0026thinsp;1.42), which was significantly higher than ChatGPT-4o (95.43\u0026thinsp;\u0026plusmn;\u0026thinsp;1.98, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and ChatGPT-5 (95.19\u0026thinsp;\u0026plusmn;\u0026thinsp;2.27, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). No statistically significant difference was observed between ChatGPT-4o and ChatGPT-5 in the text-only component (p\u0026thinsp;=\u0026thinsp;1.00).\u003c/p\u003e\n\u003ch3\u003ePerformance on Image-based Examination\u003c/h3\u003e\n\u003cp\u003eIn the image-based examination, performance differed significantly among groups (Kruskal\u0026ndash;Wallis test, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Gemini 3 Pro achieved the highest mean score (93.94\u0026thinsp;\u0026plusmn;\u0026thinsp;3.41), followed by ChatGPT-5 (87.24\u0026thinsp;\u0026plusmn;\u0026thinsp;6.97). Both models scored higher than the medical student cohort (82.94\u0026thinsp;\u0026plusmn;\u0026thinsp;9.69), with statistically significant differences (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 and p\u0026thinsp;=\u0026thinsp;0.015, respectively).\u003c/p\u003e \u003cp\u003eMedical students achieved higher mean image-based scores than ChatGPT-4o (77.36\u0026thinsp;\u0026plusmn;\u0026thinsp;5.40) and Gemini 2.5 Flash (70.34\u0026thinsp;\u0026plusmn;\u0026thinsp;5.08). These differences were statistically significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for both comparisons). Gemini 2.5 Flash had the lowest mean score among all groups in the image-based component (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB).\u003c/p\u003e\n\u003ch3\u003eAnalysis by Text-only Question Type\u003c/h3\u003e\n\u003cp\u003eText-only performance was further analyzed according to question format (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). In multiple-choice questions, Gemini 3 Pro achieved the highest accuracy (98.5%), followed by Gemini 2.5 Flash (97.9%), ChatGPT-5 (95.6%), and ChatGPT-4o (95.1%). Medical students achieved a mean accuracy of 82.0%. Differences between students and all VLMs were statistically significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy of vision\u0026ndash;language models and medical students across different text-only question formats.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGroup\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMultiple-choice\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMultiple-selection\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMatching\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGemini 3 Pro\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e98.52\u0026thinsp;\u0026plusmn;\u0026thinsp;1.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e98.31\u0026thinsp;\u0026plusmn;\u0026thinsp;2.35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e98.40\u0026thinsp;\u0026plusmn;\u0026thinsp;3.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChatGPT-5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e95.94\u0026thinsp;\u0026plusmn;\u0026thinsp;3.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e96.41\u0026thinsp;\u0026plusmn;\u0026thinsp;3.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e88.67\u0026thinsp;\u0026plusmn;\u0026thinsp;6.77\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChatGPT-4o\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e96.77\u0026thinsp;\u0026plusmn;\u0026thinsp;2.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e94.92\u0026thinsp;\u0026plusmn;\u0026thinsp;4.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e89.60\u0026thinsp;\u0026plusmn;\u0026thinsp;5.82\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGemini 2.5 Flash\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e96.26\u0026thinsp;\u0026plusmn;\u0026thinsp;2.29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e97.23\u0026thinsp;\u0026plusmn;\u0026thinsp;2.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e95.73\u0026thinsp;\u0026plusmn;\u0026thinsp;4.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMedical Students\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e81.99\u0026thinsp;\u0026plusmn;\u0026thinsp;10.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e89.45\u0026thinsp;\u0026plusmn;\u0026thinsp;8.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e90.19\u0026thinsp;\u0026plusmn;\u0026thinsp;10.35\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003cb\u003eFootnote\u003c/b\u003e: Data are presented as mean accuracy percentages\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn multiple-selection questions, Gemini 3 Pro (98.3%) and Gemini 2.5 Flash (97.2%) achieved the highest accuracies, followed by ChatGPT-4o (94.8%) and ChatGPT-5 (94.1%). Medical students achieved a mean accuracy of 85.6%. All VLMs scored significantly higher than students (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003eIn matching questions, Gemini 3 Pro achieved the highest accuracy (98.4%), followed by Gemini 2.5 Flash (95.7%). Medical students achieved a mean accuracy of 90.2%, which was higher than ChatGPT-4o (89.6%) and ChatGPT-5 (88.7%). Differences between students and both ChatGPT models were statistically significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05).\u003c/p\u003e\n\u003ch3\u003eAnalysis by Image-based Sub-domains\u003c/h3\u003e\n\u003cp\u003ePerformance in image-based questions was analyzed across seven predefined sub-domains (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Gemini 3 Pro achieved the highest accuracy across all sub-domains, particularly in descriptive reporting (99.4%) and diagnosis (96.7%).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracy of vision\u0026ndash;language models across distinct image-based dermatological sub-domains.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSub-domain\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGemini 3 Pro\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eChatGPT-4o\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eGemini 2.5 Flash\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eElementary lesion recognition\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e89.40\u0026thinsp;\u0026plusmn;\u0026thinsp;7.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e79.90\u0026thinsp;\u0026plusmn;\u0026thinsp;12.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e79.60\u0026thinsp;\u0026plusmn;\u0026thinsp;16.44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e75.00\u0026thinsp;\u0026plusmn;\u0026thinsp;15.75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDescriptive reporting\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e99.40\u0026thinsp;\u0026plusmn;\u0026thinsp;3.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e99.00\u0026thinsp;\u0026plusmn;\u0026thinsp;3.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e96.60\u0026thinsp;\u0026plusmn;\u0026thinsp;8.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e77.60\u0026thinsp;\u0026plusmn;\u0026thinsp;18.58\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDiagnosis\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e96.67\u0026thinsp;\u0026plusmn;\u0026thinsp;10.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e96.67\u0026thinsp;\u0026plusmn;\u0026thinsp;8.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e90.00\u0026thinsp;\u0026plusmn;\u0026thinsp;14.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e86.40\u0026thinsp;\u0026plusmn;\u0026thinsp;16.38\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eEtiology\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e95.00\u0026thinsp;\u0026plusmn;\u0026thinsp;15.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e94.80\u0026thinsp;\u0026plusmn;\u0026thinsp;15.15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e61.80\u0026thinsp;\u0026plusmn;\u0026thinsp;28.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e53.00\u0026thinsp;\u0026plusmn;\u0026thinsp;35.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDifferential diagnosis\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e87.80\u0026thinsp;\u0026plusmn;\u0026thinsp;13.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e73.40\u0026thinsp;\u0026plusmn;\u0026thinsp;23.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e62.67\u0026thinsp;\u0026plusmn;\u0026thinsp;26.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e54.73\u0026thinsp;\u0026plusmn;\u0026thinsp;24.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTreatment\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e88.60\u0026thinsp;\u0026plusmn;\u0026thinsp;7.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e86.20\u0026thinsp;\u0026plusmn;\u0026thinsp;12.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e79.60\u0026thinsp;\u0026plusmn;\u0026thinsp;10.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e58.60\u0026thinsp;\u0026plusmn;\u0026thinsp;23.04\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGeneral knowledge\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e97.92\u0026thinsp;\u0026plusmn;\u0026thinsp;4.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c3\"\u003e \u003cp\u003e86.24\u0026thinsp;\u0026plusmn;\u0026thinsp;16.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c4\"\u003e \u003cp\u003e72.56\u0026thinsp;\u0026plusmn;\u0026thinsp;11.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c5\"\u003e \u003cp\u003e72.48\u0026thinsp;\u0026plusmn;\u0026thinsp;13.89\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eData are presented as mean accuracy percentages\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation.\u003c/p\u003e \u003cp\u003eIn the etiology sub-domain, Gemini 3 Pro (95.0%) and ChatGPT-5 (94.8%) achieved higher accuracies than ChatGPT-4o (61.8%) and Gemini 2.5 Flash (53.0%). Differences between model performances in this sub-domain were statistically significant (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001).\u003c/p\u003e \u003cp\u003eIn treatment-related questions, Gemini 3 Pro and ChatGPT-5 again achieved higher accuracies compared with ChatGPT-4o, Gemini 2.5 Flash, and the student cohort. Performance patterns across remaining sub-domains are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eComparison of Text-only and Image-based Performance\u003c/h2\u003e \u003cp\u003eWithin-group comparisons between text-only and image-based scores were performed using the Wilcoxon signed-rank test (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). All groups demonstrated statistically significant differences between modalities (p\u0026thinsp;\u0026lt;\u0026thinsp;0.05).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe largest mean score differences between text-only and image-based components were observed for Gemini 2.5 Flash (26.1 points, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and ChatGPT-4o (18.0 points, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). ChatGPT-5 demonstrated a smaller difference of 8.0 points (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), while Gemini 3 Pro showed the smallest difference among models (4.5 points, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Notably, medical students demonstrated the most balanced profile with a mean difference of 1.97 points between modalities (p\u0026thinsp;=\u0026thinsp;0.002).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCorrelation with Examination Difficulty\u003c/h3\u003e\n\u003cp\u003eTo assess the relationship between examination difficulty and group performance, mean scores of each VLM were correlated with mean student scores across the 10 exam sessions (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn the text-only component, no significant correlation was observed between student scores and VLM scores. In contrast, in the image-based component, positive correlations were observed between student performance and the performance of ChatGPT-4o and Gemini 2.5 Flash, whereas Gemini 3 Pro maintained consistently high performance across sessions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eModel Output Variability\u003c/h3\u003e\n\u003cp\u003eModel output variability was assessed by analyzing score distributions across five repeated runs for each exam session (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). Gemini 3 Pro demonstrated the lowest mean standard deviation across repeated runs in both text-only (SD\u0026thinsp;=\u0026thinsp;0.72) and image-based (SD\u0026thinsp;=\u0026thinsp;1.70) components, while higher variability was observed for Gemini 2.5 Flash, particularly in the image-based component (SD\u0026thinsp;=\u0026thinsp;3.23). ChatGPT-4o and ChatGPT-5 demonstrated intermediate levels of variability.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThis study provides a structured comparison of contemporary VLMs and medical students within an authentic undergraduate dermatology assessment framework, revealing clear modality-dependent differences in performance. By analyzing both text-only and image-based examination components across multiple exam sessions, we found that strong performance in text-based dermatologic assessments does not consistently translate to image-based reasoning. While several VLMs achieved near-ceiling scores on text-only tasks, performance diverged substantially on image-based questions, with marked variability across model architectures. These findings indicate that multimodal competence in dermatology remains heterogeneous and task-dependent, underscoring the importance of evaluating visual and textual reasoning separately rather than inferring overall capability from text-based performance alone.\u003c/p\u003e \u003cp\u003eIn the text-only domain, all evaluated VLMs achieved near-ceiling performance, exceeding the mean scores of the medical student cohort across exam sessions. Moreover, model performance remained largely invariant to fluctuations in exam difficulty that significantly affected student scores. These findings are consistent with prior studies demonstrating that LLMs can reach or exceed passing thresholds on national licensing and specialty board examinations, reflecting strong capabilities in factual recall and structured medical knowledge application.\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e,\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e However, aggregate text-based scores obscure meaningful differences between question formats. While VLMs consistently outperformed students on multiple-choice and multiple-select items, students achieved comparable or slightly higher accuracy than ChatGPT-4o and ChatGPT-5 on matching-type questions, suggesting that certain associative reasoning tasks without contextual scaffolding may still favor human cognitive strategies.\u003c/p\u003e \u003cp\u003eThe image-based examination component revealed a markedly different performance hierarchy. While Gemini 3 Pro and ChatGPT-5 achieved higher mean scores than medical students, students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. This reversal underscores that visual diagnostic competence remains a challenging domain for some VLMs, despite their strong performance on text-based tasks. Similar modality-dependent performance patterns have been reported in recent evaluations, where accuracy declined for items incorporating images or complex visual contexts.\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e,\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e Notably, the student cohort exhibited the smallest performance difference between text-only and image-based tasks, indicating a relatively consistent integration of theoretical knowledge and visual interpretation. In contrast, earlier or lighter model variants showed substantial performance declines when transitioning from text to image-based tasks, reflecting limited multimodal consistency. Consequently, text-based accuracy alone is an incomplete proxy for multimodal competence and should be interpreted cautiously when evaluating the readiness of VLMs for visually intensive clinical domains such as dermatology.\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eSub-domain analyses further clarify the nature of this discrepancy. While most models performed well in tasks emphasizing visual description and pattern recognition, notable limitations emerged in domains requiring etiological reasoning from visual input. For example, ChatGPT-4o and Gemini 2.5 Flash demonstrated reduced accuracy in etiology-related questions despite achieving moderate diagnostic accuracy. This pattern suggests that correct visual labeling does not necessarily coincide with coherent pathophysiological reasoning, a distinction that has also been emphasized in prior analyses of AI-assisted dermatologic diagnosis.\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e In contrast, Gemini 3 Pro and ChatGPT-5 maintained higher performance across etiological and treatment-related domains, indicating improved integration of visual features with underlying clinical knowledge.\u003c/p\u003e \u003cp\u003eModel reliability represents an additional and clinically relevant dimension. Repeated exam administrations revealed substantial variability in output stability across models. Gemini 3 Pro exhibited the lowest performance variance across both modalities, whereas Gemini 2.5 Flash showed higher stochastic variability, particularly in image-based tasks. Importantly, such variability may limit the interpretability and pedagogical reliability of model outputs in educational settings. These findings highlight that average accuracy alone is insufficient for evaluating AI systems intended for educational or clinical support, as reproducibility and consistency are essential prerequisites for trust and safe deployment.\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eSeveral methodological strengths enhance the relevance of this study. The use of original, non-English examination materials provides insight into multilingual model performance in real educational settings. The inclusion of a large human cohort and multiple repeated AI evaluations allows robust statistical comparison and characterization of performance variability. Nevertheless, important limitations should be acknowledged. This was a single-center study reflecting one institutional curriculum, and the examinations, while clinically oriented, remain structured academic assessments. Static images cannot fully capture the complexity of real-world dermatologic evaluation, which often incorporates dermoscopy, palpation, and longitudinal clinical context.\u003c/p\u003e \u003cp\u003eTaken together, these findings indicate that while contemporary VLMs achieve very high performance on text-based dermatologic assessments, their image-based diagnostic capabilities remain heterogeneous and strongly model-dependent. Medical students, although achieving lower absolute scores in text-only tasks, display greater consistency across modalities, reflecting the integrative nature of clinical training. Continued improvements in multimodal reasoning are evident in newer model generations; however, careful validation and human oversight remain essential. Rather than replacing human expertise, VLMs may best be positioned as complementary tools in dermatology education, supporting learning and assessment within a controlled, human-in-the-loop framework.\u003c/p\u003e "},{"header":"METHODS","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003eStudy Design and Setting\u003c/h2\u003e \u003cp\u003eThis comparative cross-sectional study assessed the performance of VLMs and fifth-year medical students on standardized dermatology clerkship examinations at Istanbul Medeniyet University Faculty of Medicine. The dataset comprised ten consecutive clerkship rotations conducted between September 2023 and January 2025. Both the text-only and image-based examinations were mandatory components of the official clerkship assessment. All analyses were performed using the original Turkish questions and responses to evaluate performance in the language of instruction.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eParticipants\u003c/h2\u003e \u003cp\u003eA total of 252 fifth-year medical students participated across ten clerkship groups (21\u0026ndash;29 students per group; mean\u0026thinsp;\u0026asymp;\u0026thinsp;25). Because the examinations were a required part of the rotation, no exclusion criteria were applied.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eExam Structure and Scoring\u003c/h2\u003e \u003cp\u003eAcross the ten exam sessions, the dataset included 500 text-only items and 200 image-based items. Each session consisted of:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eText-only examination (50 items): 31 single-answer multiple-choice (Q1\u0026ndash;31), 13 multiple-select (Q32\u0026ndash;44), and 6 matching questions (Q45\u0026ndash;50).\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eImage-based examination (20 items): structured open-ended questions spanning seven sub-domains: elementary lesion recognition, descriptive reporting, diagnosis, etiology, differential diagnosis, treatment, and general dermatological knowledge.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eText-only items were graded using predefined answer keys, whereas image-based items were graded using weighted rubrics. The overall examination score was calculated as a weighted combination of modalities (text-only 60%, image-based 40%). The detailed exam structure and scoring framework is summarized in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eStructure and scoring scheme of the dermatology clerkship examinations.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExam component\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSubcategory (question range)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo. of items\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eScoring scheme\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMax score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003eText-only exam\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMultiple-choice (Q1\u0026ndash;31)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2 points per correct answer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e62\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMultiple-select (Q32\u0026ndash;44)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e(Correct selections\u0026thinsp;\u0026divide;\u0026thinsp;3) \u0026times; 2; max 2 points per item; invalid (0 points) if\u0026thinsp;\u0026gt;\u0026thinsp;3 choices selected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMatching (Q45\u0026ndash;50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e(Correct matches\u0026thinsp;\u0026divide;\u0026thinsp;5) \u0026times; 2; max 2 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e50\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e100\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"7\" rowspan=\"8\"\u003e \u003cp\u003e\u003cb\u003eImage-based exam\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eElementary lesion recognition (Q1\u0026ndash;4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWeighted rubric; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescriptive reporting (Q5\u0026ndash;6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWeighted rubric; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDiagnosis (Q7\u0026ndash;9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSingle-answer; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEtiology (Q10\u0026ndash;11)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSingle-answer; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDifferential diagnosis (Q12\u0026ndash;13)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWeighted rubric; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTreatment (Q14\u0026ndash;15)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWeighted rubric or single-answer; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGeneral dermatological knowledge (Q16\u0026ndash;20)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWeighted rubric or single-answer; max 5 points per item\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e20\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e100\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"5\"\u003e\u003cb\u003eFootnote\u003c/b\u003e: Text-only and image-based components were each scaled to a maximum of 100 points. For multiple-select items, partial credit was calculated as (number of correct selections\u0026thinsp;\u0026divide;\u0026thinsp;3) \u0026times; 2 (maximum 2 points per item); responses selecting more than three options were scored as 0. For matching items, partial credit was calculated as (number of correct matches\u0026thinsp;\u0026divide;\u0026thinsp;5) \u0026times; 2 (maximum 2 points per item). \u0026lsquo;Weighted rubric\u0026rsquo; indicates predefined point allocations (up to 5 points per item) based on faculty-approved grading criteria (e.g., awarding credit for specific expected elements within an open-ended response). Overall score was computed as a weighted average: 0.60 \u0026times; text-only score\u0026thinsp;+\u0026thinsp;0.40 \u0026times; image-based score.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eAI Models and Testing Procedure\u003c/h2\u003e \u003cp\u003eFour VLMs developed by OpenAI (ChatGPT-4o, ChatGPT-5) and Google (Gemini 2.5 Flash, Gemini 3 Pro) were evaluated between May and October 2025. Model testing used standardized, format-specific Turkish prompts applied identically across all models and examination items. Text-only items were submitted in grouped batches with explicit formatting instructions (multiple-choice, multiple-select, or matching). Image-based items were submitted individually and required structured open-ended responses with predefined constraints on the type and number of expected answers. A representative testing workflow is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. To account for stochastic variability in generative outputs, each exam session was repeated five times per model.\u003c/p\u003e \u003cp\u003eFor transparency, standardized prompt templates and representative examples are provided in Supplementary Information I, and the full content of the first evaluated examination (22 September 2023) with English translations is provided in Supplementary Information II\u0026ndash;III, together with the corresponding expert-validated answer keys.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eGold-standard Answers and Evaluation of Model Outputs\u003c/h2\u003e \u003cp\u003eFaculty members from relevant subspecialty areas authored the examination items. All questions were compiled, reviewed, and approved by the clerkship coordinator. Gold-standard answers (text-only) and grading rubrics (image-based) were prepared by the item authors and verified by the clerkship coordinator.\u003c/p\u003e \u003cp\u003e VLM responses were scored by the clerkship coordinator and a senior dermatology resident through collaborative review with task sharing, reflecting the routine grading workflow used in the clerkship. No independent parallel scoring, averaging, or formal consensus procedure was applied.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eAvailability of Student Image-based Sub-domain Data\u003c/h2\u003e \u003cp\u003eStudent image-based data were obtained retrospectively from official clerkship records. While total image-based examination scores were available for all students, sub-domain\u0026ndash;level rubric components were not recorded in the examination system at the time of assessment. Accordingly, sub-domain analyses for image-based performance could be conducted only for VLMs, for which rubric-based scoring was prospectively recorded at the item level.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses were performed in Python (version 3.10) using the \u003cem\u003epandas\u003c/em\u003e, \u003cem\u003enumpy\u003c/em\u003e, and \u003cem\u003escipy\u003c/em\u003e libraries. Data visualization was conducted using \u003cem\u003ematplotlib\u003c/em\u003e and \u003cem\u003eseaborn\u003c/em\u003e to generate violin plots, radar charts, and slope graphs. Normality of score distributions was assessed using the Shapiro\u0026ndash;Wilk test. As several variables deviated from normality, non-parametric statistical methods were applied. Overall between-group differences were evaluated using the Kruskal\u0026ndash;Wallis test, followed by pairwise Mann\u0026ndash;Whitney U tests for relevant group comparisons (students versus each model and between-model comparisons), with Bonferroni correction applied for multiple testing. To assess within-group differences between text-only and image-based performance (the modality gap), the Wilcoxon signed-rank test was used. Correlations between mean student scores and model scores across examination sessions were examined using Spearman\u0026rsquo;s rank correlation coefficient. Model output variability was quantified descriptively using the standard deviation of total scores across five repeated runs per model and examination session. Statistical significance was defined as a two-sided \u003cem\u003ep\u003c/em\u003e value\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u0026nbsp;information:\u003c/strong\u003e This article has no funding source.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest:\u003c/strong\u003e The authors have no conflict of interest to declare.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Approval:\u003c/strong\u003e Reviewed and approved by the Ethics Committee of Prof. Dr. Suleyman Yalcin City Hospital (Approval No.: 2025/0263) and Deanship of the Faculty of Medicine, Istanbul Medeniyet University (Approval No.: E-28298836-100-2500069724). This study was conducted retrospectively, and the requirement for informed consent for participants was waived by the Ethics Committee and confirmed by the Faculty of Medicine Deanship.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics statement:\u0026nbsp;\u003c/strong\u003eThe study was conducted in accordance with the principles of the Declaration of Helsinki.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability statement\u003c/strong\u003e: The data that support the findings of this study are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contribution:\u0026nbsp;\u003c/strong\u003eO.E. and A.Y. conceptualized and designed the study and developed the methodology. Formal analysis was performed by O.E., A.S.Ş., and A.Y. Investigation was carried out by O.E. and A.S.Ş. Resources were provided by O.E., V.A.E., M.A.K., and M.S.G. O.E. prepared the visualizations. The original draft of the manuscript was written by O.E. and A.Y. Manuscript review and editing were performed by E.G. and B.B.D. M.S.G. supervised the study. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgement:\u0026nbsp;\u003c/strong\u003eAbdurrahim Yilmaz has been funded by the President\u0026rsquo;s PhD Scholarships at Imperial College London, which solely supported his academic studies.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eEsteva, A. et al. Dermatologist-level classification of skin cancer with deep neural networks. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e542\u003c/b\u003e, 115\u0026ndash;118 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTopol, E. J. High-performance medicine: the convergence of human and artificial intelligence. \u003cem\u003eNat. Med.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e, 44\u0026ndash;56 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGordon, M. et al. A scoping review of artificial intelligence in medical education: BEME Guide 84. \u003cem\u003eMed. Teach.\u003c/em\u003e \u003cb\u003e46\u003c/b\u003e, 446\u0026ndash;470 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVrdoljak, J., Boban, Z., Vilović, M., Kumrić, M. \u0026amp; Božić, J. A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. \u003cem\u003eHealthcare\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 603 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThirunavukarasu, A. J. et al. Large language models in medicine. \u003cem\u003eNat. Med.\u003c/em\u003e \u003cb\u003e29\u003c/b\u003e, 1930\u0026ndash;1940 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinghal, K. et al. Large language models encode clinical knowledge. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e620\u003c/b\u003e, 172\u0026ndash;180 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoor, M. et al. Foundation models for generalist medical artificial intelligence. \u003cem\u003eNature\u003c/em\u003e \u003cb\u003e616\u003c/b\u003e, 259\u0026ndash;265 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHartsock, I. \u0026amp; Rasool, G. Vision-language models for medical report generation and visual question answering: a review. \u003cem\u003eFront. Artif. Intell.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 1430984 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrones, F., Marikkar, U., Parsons, G., Szmul, A. \u0026amp; Mahdi, A. Review of multimodal machine learning approaches in healthcare. \u003cem\u003eInf. Fusion\u003c/em\u003e. \u003cb\u003e114\u003c/b\u003e, 102690 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Z. et al. A perspective for adapting generalist AI to specialized medical AI applications and their challenges. \u003cem\u003eNpj Digit. Med.\u003c/em\u003e \u003cb\u003e8\u003c/b\u003e, 429 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan, S. et al. A multimodal vision foundation model for clinical dermatology. \u003cem\u003eNat. Med.\u003c/em\u003e \u003cb\u003e31\u003c/b\u003e, 2691\u0026ndash;2702 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, J. et al. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 5649 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYilmaz, A. et al. Resource-efficient medical vision language model for dermatology via a synthetic data generation framework. 05.17.25327785 Preprint at (2025). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2025.05.17.25327785\u003c/span\u003e\u003cspan address=\"10.1101/2025.05.17.25327785\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKung, T. H. et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. \u003cem\u003ePLOS Digit. Health\u003c/em\u003e. \u003cb\u003e2\u003c/b\u003e, e0000198 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGilson, A. et al. How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e, e45312 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrin, D. et al. Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 16492 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSadeq, M. A. et al. AI chatbots show promise but limitations on UK medical exam questions: a comparative performance study. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e, 18859 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCasals-Farre, O. et al. Assessing ChatGPT 4.0\u0026rsquo;s Capabilities in the United Kingdom Medical Licensing Examination (UKMLA): A Robust Categorical Analysis. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 13031 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, H. J. et al. Performance evaluation of large language models on Korean medical licensing examination: a three-year comparative analysis. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 36082 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJung, L. B. et al. ChatGPT Passes German State Examination in Medicine With Picture Questions Omitted. \u003cem\u003eDtsch. Arzteblatt Int.\u003c/em\u003e \u003cb\u003e120\u003c/b\u003e, 373\u0026ndash;374 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakagi, S., Watari, T., Erabi, A. \u0026amp; Sakaguchi, K. Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: Comparison Study. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e, e48002 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRojas, M., Rojas, M., Burgess, V., Toro-P\u0026eacute;rez, J. \u0026amp; Salehi, S. Exploring the Performance of ChatGPT Versions 3.5, 4, and 4 With Vision in the Chilean Medical Licensing Examination: Observational Study. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, e55048\u0026ndash;e55048 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiyazaki, Y. et al. Performance of ChatGPT-4o on the Japanese Medical Licensing Examination: Evalution of Accuracy in Text-Only and Image-Based Questions. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, e63129\u0026ndash;e63129 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo, D. et al. Evaluating the performance of GPT-3.5, GPT-4, and GPT-4o in the Chinese National Medical Licensing Examination. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 14119 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuerra, G. A. et al. GPT-4 Artificial Intelligence Model Outperforms ChatGPT, Medical Students, and Neurosurgery Residents on Neurosurgery Written Board-Like Questions. \u003cem\u003eWorld Neurosurg.\u003c/em\u003e \u003cb\u003e179\u003c/b\u003e, e160\u0026ndash;e165 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMassey, P. A., Montgomery, C. \u0026amp; Zhang, A. S. Comparison of ChatGPT\u0026ndash;3.5, ChatGPT-4, and Orthopaedic Resident Performance on Orthopaedic Assessment Examinations. \u003cem\u003eJ. Am. Acad. Orthop. Surg.\u003c/em\u003e \u003cb\u003e31\u003c/b\u003e, 1173\u0026ndash;1179 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoos, J., Kasapovic, A., Jansen, T. \u0026amp; Kaczmarczyk, R. Artificial Intelligence in Medical Education: Comparative Analysis of ChatGPT, Bing, and Medical Students in Germany. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e, e46482 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeyer, A., Riese, J. \u0026amp; Streichert, T. Comparison of the Performance of GPT-3.5 and GPT-4 With That of Medical Students on the Written German Medical Licensing Examination: Observational Study. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, e50965 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBahir, D. et al. Gemini AI vs. ChatGPT: A comprehensive examination alongside ophthalmology residents in medical knowledge. \u003cem\u003eGraefes Arch. Clin. Exp. Ophthalmol.\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00417-024-06625-4\u003c/span\u003e\u003cspan address=\"10.1007/s00417-024-06625-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZengin, A., Ulfanov, O., Bag, Y. M. \u0026amp; Ulas, M. Artificial Intelligence Versus Medical Students in General Surgery Exam. \u003cem\u003eIndian J. Surg.\u003c/em\u003e \u003cb\u003e87\u003c/b\u003e, 68\u0026ndash;73 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, X. \u0026amp; Chen, W. The performance of ChatGPT on medical image-based assessments and implications for medical education. \u003cem\u003eBMC Med. Educ.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e, 1192 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBehrmann, J. et al. Chat generative pre-trained transformer\u0026rsquo;s performance on dermatology-specific questions and its implications in medical education. \u003cem\u003eJ. Med. Artif. Intell.\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, 16\u0026ndash;16 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePassby, L., Jenko, N. \u0026amp; Wernham, A. Performance of ChatGPT on Specialty Certificate Examination in Dermatology multiple-choice questions. \u003cem\u003eClin. Exp. Dermatol.\u003c/em\u003e \u003cb\u003e49\u003c/b\u003e, 722\u0026ndash;727 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan, K. S. \u0026amp; Fan, K. H. Dermatological Knowledge and Image Analysis Performance of Large Language Models Based on Specialty Certificate Examination in Dermatology. \u003cem\u003eDermato\u003c/em\u003e \u003cb\u003e4\u003c/b\u003e, 124\u0026ndash;135 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eG\u0026ouml;\u0026ccedil;er G\u0026uuml;rok, N. \u0026amp; \u0026Ouml;zt\u0026uuml;rk, S. The Performance of AI in Dermatology Exams: The Exam Success and Limits of ChatGPT. \u003cem\u003eJ. Cosmet. Dermatol.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e, e70244 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAtılan, A. U. \u0026amp; \u0026Ccedil;etin, N. Benchmarking Large Language Models on the Turkish Dermatology Board Exam: A Comparative Multilingual Analysis. \u003cem\u003eTurk. J. Dermatol.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 126\u0026ndash;133 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBicknell, B. T. et al. ChatGPT-4 Omni Performance in USMLE Disciplines and Clinical Skills: Comparative Analysis. \u003cem\u003eJMIR Med. Educ.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, e63430 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrin, D. et al. How GPT models perform on the United States medical licensing examination: a systematic review. \u003cem\u003eDiscov Appl. Sci.\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, 500 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, M. et al. Evaluating the Effectiveness of advanced large language models in medical Knowledge: A Comparative study using Japanese national medical examination. \u003cem\u003eInt. J. Med. Inf.\u003c/em\u003e \u003cb\u003e193\u003c/b\u003e, 105673 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin, Q. et al. Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine. \u003cem\u003eNpj Digit. Med.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, 190 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, Z., Xin, C., Yu, Y., Xia, J. \u0026amp; Han, L. AI dermatology: Reviewing the frontiers of skin cancer detection technologies. \u003cem\u003eIntell. Oncol.\u003c/em\u003e \u003cb\u003e1\u003c/b\u003e, 89\u0026ndash;104 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan, H. Challenges of reproducible AI in biomedical data science. \u003cem\u003eBMC Med. Genomics\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e, 8, (2025). s12920-024-02072\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Vision-Language Models, Medical Education, Dermatology, Medical Students, Medical Examinations, Large Language Models","lastPublishedDoi":"10.21203/rs.3.rs-8480126/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8480126/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eVision\u0026ndash;language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs (ChatGPT-4o, ChatGPT-5, Gemini 2.5 Flash, and Gemini 3 Pro) with fifth-year medical students on ten consecutive dermatology clerkship examinations administered between September 2023 and January 2025. Examinations combined text-only questions (multiple-choice, multiple-select, and matching; 60% of total score) with image-based, structured open-ended questions (40%) spanning seven dermatologic sub-domains. Model outputs were evaluated using expert-validated answer keys and grading rubrics, with repeated runs to assess output variability. All VLMs significantly outperformed medical students on text-only examinations (mean scores\u0026thinsp;\u0026gt;\u0026thinsp;95 vs. 84.9; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), showing minimal sensitivity to exam difficulty. In contrast, image-based performance was heterogeneous: Gemini 3 Pro and ChatGPT-5 achieved higher scores than students, whereas students significantly outperformed ChatGPT-4o and Gemini 2.5 Flash. Medical students demonstrated the smallest performance gap between text-only and image-based components, indicating greater cross-modal consistency. Sub-domain analyses revealed that some models achieved accurate visual description and diagnosis but showed reduced performance in etiological and treatment reasoning. Gemini 3 Pro exhibited the highest overall accuracy and the lowest output variability across repeated evaluations. These findings indicate that while VLMs excel in text-based dermatologic assessment, multimodal competence remains uneven and model-dependent, supporting their use as complementary rather than standalone tools in dermatology education.\u003c/p\u003e","manuscriptTitle":"Performance of Vision–Language Models Compared with 252 Medical Students on Text-only and Image-based Dermatology Examinations","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-09 16:45:10","doi":"10.21203/rs.3.rs-8480126/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-04-30T13:49:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"245401191643568529342021827521538770535","date":"2026-04-27T09:03:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"19398823240616755151696147676990405474","date":"2026-04-24T06:36:55+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-22T17:43:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"240849804107760919955562239317285064873","date":"2026-04-13T08:15:35+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-05T15:06:54+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-31T17:44:27+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-01-05T20:28:52+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-02T16:34:59+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-01-02T16:26:48+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"27bf481e-79ab-41a3-98eb-1726a85ce9c3","owner":[],"postedDate":"April 9th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-04-30T13:49:12+00:00","index":107,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":65985256,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":65985257,"name":"Health sciences/Diseases"},{"id":65985258,"name":"Health sciences/Health care"},{"id":65985259,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-04-09T16:45:10+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-09 16:45:10","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8480126","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8480126","identity":"rs-8480126","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0