More Harm than Help? Evaluating the Capabilities of Vision-Language Models in Neurological Image Analysis

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Objectives: This study evaluates the performance of both open-source and commercial Vision Language Models (VLMs) in interpreting radiological images of neurological diseases, comparing their diagnostic accuracy to that of experienced neuroradiologists. Methods: A dataset comprising 100 cases of brain and spine pathologies with confirmed diagnoses was curated from the Radiopaedia database to reflect routine clinical neuroradiology practice. Five neuroradiologists reviewed the cases—including imaging and case presentations—to determine the most probable diagnosis. In parallel, five VLMs (Gemini 2.0, GPT-4o1-Preview, Llama 3.2 90b, Qwen 2.5, and Grok-2-vision) were provided with the same cases and tasked with generating three differential diagnoses along with their reasoning. Two neuroradiologists then evaluated the accuracy of both the single most probable diagnosis and the top three diagnoses produced by the VLMs, as well as the rationale provided, and assessed the potential for harmful outcomes based on the VLM outputs. Results: Neuroradiologists achieved a mean diagnostic accuracy of 86.2%, significantly outperforming all VLMs. Among the models, Gemini 2.0 achieved the highest accuracy at 35% with 28% of its diagnoses deemed potentially harmful, while Grok-2-vision had the lowest accuracy at 9% with 45% of its outputs categorized as harmful. All models demonstrated a trend toward slightly lower accuracy with an increasing number of images, however the strength of this relationship was modest. Evaluation of potential harm revealed that treatment delay was the most common risk for VLMs, ranging between 28% for Gemini 2.0 and 45% for Grok-2-vision. Error analysis indicated that the most frequent causes of misdiagnosis were incorrect anatomic classification—with error rates ranging from 26% for Gemini 2.0 to 53% for Grok-2-vision —and inaccurate description of imaging findings, which ranged from 35% for Gemini 2.0 to 72% for Grok-2-vision. Conclusion: While VLMs hold promise for enhancing radiological workflows, the current state-of-the-art of open-source and commercial models is far from being reliable for the interpretation of radiological images of neurological diseases.
Full text 96,635 characters · extracted from preprint-html · click to expand
More Harm than Help? Evaluating the Capabilities of Vision-Language Models in Neurological Image Analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article More Harm than Help? Evaluating the Capabilities of Vision-Language Models in Neurological Image Analysis Aymen Meddeb, Ida Rangus, Paolo Pagano, Insaf Dekhil, Soumaya Jelassi, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6183659/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 17 Nov, 2025 Read the published version in npj Digital Medicine → Version 1 posted 9 You are reading this latest preprint version Abstract Objectives: This study evaluates the performance of both open-source and commercial Vision Language Models (VLMs) in interpreting radiological images of neurological diseases, comparing their diagnostic accuracy to that of experienced neuroradiologists. Methods: A dataset comprising 100 cases of brain and spine pathologies with confirmed diagnoses was curated from the Radiopaedia database to reflect routine clinical neuroradiology practice. Five neuroradiologists reviewed the cases—including imaging and case presentations—to determine the most probable diagnosis. In parallel, five VLMs (Gemini 2.0, GPT-4o1-Preview, Llama 3.2 90b, Qwen 2.5, and Grok-2-vision) were provided with the same cases and tasked with generating three differential diagnoses along with their reasoning. Two neuroradiologists then evaluated the accuracy of both the single most probable diagnosis and the top three diagnoses produced by the VLMs, as well as the rationale provided, and assessed the potential for harmful outcomes based on the VLM outputs. Results: Neuroradiologists achieved a mean diagnostic accuracy of 86.2%, significantly outperforming all VLMs. Among the models, Gemini 2.0 achieved the highest accuracy at 35% with 28% of its diagnoses deemed potentially harmful, while Grok-2-vision had the lowest accuracy at 9% with 45% of its outputs categorized as harmful. All models demonstrated a trend toward slightly lower accuracy with an increasing number of images, however the strength of this relationship was modest. Evaluation of potential harm revealed that treatment delay was the most common risk for VLMs, ranging between 28% for Gemini 2.0 and 45% for Grok-2-vision. Error analysis indicated that the most frequent causes of misdiagnosis were incorrect anatomic classification—with error rates ranging from 26% for Gemini 2.0 to 53% for Grok-2-vision —and inaccurate description of imaging findings, which ranged from 35% for Gemini 2.0 to 72% for Grok-2-vision. Conclusion: While VLMs hold promise for enhancing radiological workflows, the current state-of-the-art of open-source and commercial models is far from being reliable for the interpretation of radiological images of neurological diseases. Health sciences/Medical research Health sciences/Neurology Figures Figure 1 Figure 2 Figure 3 Introduction The development of Large Language Models (LLMs) has expanded the potential applications of artificial intelligence (AI) in healthcare, influencing various aspects of medical practice[ 1 – 3 ]. Models such as OpenAI’s GPT, Google’s Gemini, or Meta’s Llama have demonstrated substantial potential for improving the efficiency and quality of healthcare delivery. Studies indicate that LLMs can extract information from electronic health records[ 1 , 4 ], summarize[ 5 ], translate[ 6 ] and structure medical texts[ 7 , 8 ], as well as answer medical board examination questions[ 9 , 10 ]. Furthermore, these models have been evaluated for their ability to analyze complex medical cases, suggest differential diagnoses[ 11 ] and support clinical decision-making processes[ 12 , 13 ] . Unlike LLMs, which are exclusively trained on textual data, Vision Language Models (VLMs) leverage multimodal datasets that pair visual information, such as images or videos, with corresponding text annotations[ 14 , 15 ]. The dual training enables VLMs to create joint embeddings that capture the relationships between visual and textual modalities. By combining these modalities, VLMs are applicable to tasks that require both image interpretation and contextual understanding, which are key components of radiological assessment [ 16 ] . While initial studies have primarily assessed the diagnostic performance of VLMs in closed medical visual question-answering tasks [ 17 ], critical aspects such as the quality of generated differential diagnoses and the underlying reasoning remain underexplored[ 18 ]. This study aims to evaluate the diagnostic accuracy of open-source and commercial VLMs in diagnosing a range neuroradiological diseases across various imaging modalities, comparing their performance with expert neuroradiologists. Additionally, we assess the quality of the generated differential diagnoses and perform a comprehensive error analysis, providing insights into the maturity of VLMs and their potential for clinical integration. Methods This retrospective study did use publicly available data and was therefore exempt from institutional review board approval. The inclusion of Radiopaedia cases was approved by the licensing committee under the terms of “Non-commercial use of Radiopaedia content for Machine Learning.” Data Selection and Case Preparation A search for neuroradiological cases in the Radiopedia database (radiopedia.org) was conducted by an experienced neuroradiologist (A.M.) in December 2024, selecting cases that reflect the incidence of neuroradiological diseases in clinical practice. A total of 100 cases were included based on the following criteria: 1) adequate image quality, 2) availability of a case presentation, and 3) confirmed diagnosis. Cases were stratified by complexity into three categories: straightforward, intermediate, and challenging. For each case, between 1 and 4 representative images were chosen, and all images were downloaded in JPG format. LLM Inference Three proprietary LLMs—Google’s Gemini 2.0, OpenAI’s GPT-4o1-Preview, and XAI’s Grok-2-vision —and two open-source LLMs—Meta’s Llama 3.2 90b and Alibaba’s Qwen—were selected to analyze the cases. Models’ temperature was set at 1, to ensure consistency. Access to these LLMs was obtained via API keys and Inference for all models was made between January 14th and February 4th. A standardized base prompt was used for all models and reads as follows: You are an expert neuroradiologist. Read the case presentation: {presentation}. Analyze the images and provide the modality, a description of the images, the most probable diagnosis, and the two most important differential diagnoses. The source code for the methodology used in this study is publicly available on GitHub ( https://github.com/Meddebma/VLM_Neurorad_Benchmark ). Neuroradiologist Rating Five neuroradiologists, with professional experience ranging from 8 to 17 years, were presented with the cases, including images and clinical histories, via a Google Forms questionnaire. They were instructed to provide the most probable differential diagnosis for each case. Diagnostic Accuracy Evaluation We evaluated the accuracy of the most probable diagnoses, as well as the accuracy of all three generated diagnoses. Responses from both the LLMs and neuroradiologists were evaluated based on their accuracy in identifying the exact pathological entity. Responses were rated as incorrect if they did not match the reference diagnosis precisely (e.g., “familial cavernomatosis” was rated as incorrect in a case of “amyloid angiopathy,” despite both showing susceptibility artifacts on susceptibility-weighted imaging). However, responses providing a correct but less specific diagnosis or a more granular diagnosis than the predefined reference diagnosis were considered correct (e.g., “acute ischemic stroke” was deemed correct for a case of “ischemic stroke in the territory of the right MCA”). Furthermore, we conducted a quantitative analysis to evaluate the relationship between the number of images provided per case and the diagnostic accuracy of VLMs. Harmful Diagnosis Evaluation To assess the potential harm of the most probable diagnoses generated by the models, we established strict evaluation criteria and performed a systematic harm assessment for each answer. A diagnosis was classified as "harmful" if it met at least one of the following criteria: (1) Treatment delay: if a critical, time-sensitive condition such as a ruptured aneurysm, acute ischemic stroke is not identified as the correct diagnosis. (2) Misclassification: If a benign lesion is misclassified as malignant, leading to undue patient anxiety and potentially unnecessary treatments, or a malignant lesion is misclassified as benign. (3) Overdiagnosis: if the generated diagnosis could lead to an unnecessary invasive procedure (e.g., brain biopsy for a condition that does not require it) or an inappropriate surgical intervention, or if the model misclassifies common anatomical variants as pathological findings. Two independent raters (A.M. and S.S.) reviewed the responses and assigned harm classifications. Any discrepancies in classification were resolved through discussion between the two raters until a consensus was reached. Error Analysis Error analysis of the VLM-generated diagnoses was performed by two neuroradiologists (A.M. and S.S.), who systematically categorized the errors into five distinct types: (1) Incorrect anatomic classification (i.e., misidentification of the anatomical region or side), (2) Inaccurate description of imaging findings, (3) Misidentification of imaging sequences, (4) Hallucinated findings (i.e., interpretation of non-existent pathologies as real findings), and (5) Overlooked pathologies (failure to identify existing abnormalities Any discrepancies in classification were resolved through discussion until a consensus was reached. Statistical Analysis Statistical Analysis Statistical analyses were conducted in RStudio ((version 4.3.2, R Foundation). Pairwise statistical comparisons to assess diagnostic accuracy among different VLMs and neuroradiologsits was conducted using McNemar’s test. Given the multiple comparisons between neuroradiologists and the five different VLMs, Holm-Bonferroni correction was applied to adjust for multiple hypothesis testing while maintaining statistical power. The inter-rater reliability of error analysis and harmful diagnosis evaluation was evaluated using Cohen’s kappa. An adjusted P-value of < .05 was considered statistically significant. For each model, we computed the Pearson correlation coefficient between the number of images and the corresponding accuracy scores using Python’s pandas library. Results Data Selection A total of 100 cases were included, categorized as follows: Neoplasms (n = 28), degenerative and demyelinating Diseases (n = 20), vascular diseases (n = 19), congenital and developmental disorders (n = 15), Infections (n = 9), traumatic disorders (n = 5), metabolic and toxic disorders (n = 4). Case complexities was classified as follows: straightforward (n = 65), intermediate (n = 25), challenging (n = 10). Population Characteristics are presented in Table 1 . Table 1 Population Characteristics: Sex Distribution, Age (Mean ± SD), and Number of Cases by Difficulty and Category Characteristic (n = 100) Value (%) Sex (mean distribution) 52 Male, 48 Female Number of pediatric cases (age < 18) 16 Adult Age (mean ± SD) 46.35 ± 17.36 years Number of cases (by difficulty level) Low difficulty 65 Moderate difficulty 25 High difficulty 10 Number of cases (by category) Tumors 28 Vascular diseases 19 Congenital and developmental disorders 15 Degenerative and demyelinating Diseases 20 Infections 9 Traumatic disorders 5 Metabolic and toxic disorders 4 Diagnostic Accuracy Performance metrics of the five VLMs (GPT-4o1-Preview, Gemini 2.0, Grok-2-vision, Qwen 2.5 and Llama 3.2 90b) compared with neuroradiologists are shown in Table 2 . The mean diagnostic accuracy of neuroradiologists was 86.2%, significantly outperforming all LLMs. Among the evaluated models, Gemini (35%) and GPT-4o1-Preview (27%) achieved the highest accuracy, followed by Llama 3.2 90b (24%), QWEN (23%), and Grok-2-vision (9%). Pairwise statistical comparisons using McNemar’s test showed that all LLMs had significantly lower accuracy than neuroradiologists (P < .001, Holm-Bonferroni adjusted). No significant difference was observed between Gemini and GPT-4o1-Preview (P = .73), while Grok-2-vision was significantly worse than all other LLMs (P < .001). Regarding the relationship between the number of images provided per case and the diagnostic accuracy, correlation analysis revealed negative relationships between the number of images and model accuracy across all models. Specifically, the Pearson correlation coefficients were as follows: GPT-4o1-Preview (-0.221), Gemini (-0.062), Qwen (-0.217), Llama 3.2 90b (-0.091), and Grok-2-vision (-0.146). Accuracies are presented in Table 2 and depicted in Figs. 1 and 2 . Table 2 Accuracies, harmful diagnosis categories as well as error types across five models Neuroradiologists (mean) Gemini 2.0 GPT-4o1-Preview Llama 3.2 90b Qwen 2.5 Grok-2-vision Accuracy (most probable diagnosis) 86,2% 35% 27% 24% 23% 9% Accuracy (Top 3 diagnoses) - 52% 47% 43% 40% 21% Harmful diagnoses (all categories) 15% 28% 37% 30% 32% 45% Treatment Delay 2% 16% 16% 14% 17% 28% Misclassification 11% 11% 21% 16% 14% 17% Overdiagnosis 2% 1% 0% 0% 1% 0% Error Analysis Incorrect anatomic classification - 26% 29% 51% 43% 53% Inaccurate description of imaging findings - 35% 43% 54% 58% 72% Misidentification of imaging modality/sequences - 6% 11% 21% 18% 27% Hallucinated findings - 17% 19% 26% 23% 42% Overlooked pathologies - 27% 25% 31% 33% 41% Harmful Diagnosis Evaluation The clinical harm rate, defined as the proportion of generated diagnoses classified as harmful, was highest for Grok (45%), followed by GPT-4o1-Preview (37%), QWEN (32%), Llama 3.2 90b (30%), and Gemini (28%). Figure 3 gives on overview of the percentage of misclassifications withing each category for each VLMs. Gemini 2.0 was the least harmful model with a score of 28%, with 16% in the first harm category, 11% in the second category, and 1% in the third category (Overdiagnosis). Although GPT-4o1-Preview had the second highest diagnostic accuracy, it ranked fourth in the clinical harm rate (37%), with 16% in the first category, and 21% in the second category. Grok had the highest harm rate (45%), with 28% in the first category and 17% in the second category. Holm-Bonferroni corrected pairwise comparisons showed that Grok produced significantly more harmful outputs than all other models (P < .001), while Gemini had the lowest harm rate but was not significantly different from Llama 3.2 90b (P = .41). In comparison, neuroradiologists had a combined harm rate of 15%, with 11% attributable to misclassification and 2% each for treatment delay and overdiagnosis. Supplementary file provides examples of cases with both correct and false diagnoses generated by the VLMs. Error Analysis The VLMs with varying performance levels exhibited different patterns of errors. For Gemini-2.0 and GPT-4o1-Preview, the most prevalent errors were Inaccurate description of imaging findings (35% and 43%, respectively) and Overlooked pathologies (27% and 25%, respectively). In contrast, lower-performing models such as Llama 3.2 90b, Qwen 2.5 demonstrated high Incorrect anatomic classification (51% and 43%, respectively) and high Inaccurate description of imaging findings (54% and 58%, respectively). Qwen had the highest error prevalence with the highest percentages of hallucinated findings (42%) and misidentification of imaging modality/sequences. Discussion Clinical Vision-Language Models (VLMs) have shown promise in addressing various medical applications. However, their evaluation has primarily relied on static, structured assessments—such as multiple-choice questions—that fail to capture the dynamic complexities encountered in real-world clinical practice. The results of this study highlight significant shortcomings of VLMs in accurately identifying neuroradiological diagnosis based on clinical case descriptions and images. While neuroradiologists achieved a mean accuracy of 86.2%, the best-performing LLM, Gemini 2.0, only reached 35%, with other models performing even worse. Similarly, a recent study by Busch et al. found comparable results, with GPT-4V accurately identifying the primary diagnosis in 47% of cases, when cases were presented with clinical history and multiple-choice questions[ 19 ]. This substantial performance gap highlights the current limitations of VLMs and suggests they are not yet reliable enough to be used as a clinical decision support system in neuroradiological practice. We assessed VLM performance at several levels. First, we examined the accuracy of the most probable diagnosis and whether the correct diagnosis was included among the top three differential diagnoses. Although all VLMs demonstrated improved accuracy when considering their top three choices, even the top-performing Gemini 2.0 only achieved 52% accuracy—well below the 86.2% accuracy of neuroradiologists (P < .001). Notably, accuracy varied across disease categories, with congenital and developmental disorders yielding accuracies below 10% for all VLMs. This finding likely reflects the limited availability of high-quality, publicly accessible pediatric neuroradiology resources. Consistent with a recent study by Suh et al., our results also suggest that the unique challenges of pediatric imaging—stemming from developmental variability and a diverse range of syndromic diseases—are particularly problematic for these models[ 20 ] . Regarding the relationship of number of images presented to the VLMs and accuracy rates, all correlations were negative, which imply that there might be a slight degradation in performance when more images are used, but the impact is modest (mostly between − 0.06 and − 0.22). This means that other factors than image quantity may play a more significant role in influencing model accuracy. Our accuracy evaluation further revealed that many generated diagnoses were far from the ground truth. This misalignment raises significant concerns, as reliance on VLM-generated recommendations could lead to diagnostic errors and potential harm to patients. To quantify potential risks, we defined three harm categories: (1) treatment delay, (2) misclassification, and (3) overdiagnosis. For VLMs, treatment delay was the most frequent issue, with rates ranging from 14% for Llama 3.2 90b to 28% for Grok. Misclassification was also common, occurring in 11% of cases for Gemini 2.0 and up to 17% for GPT-4o1-Preview. Overdiagnosis was rare, with only one case each for Gemini and Qwen. In contrast, harmful diagnoses proposed by neuroradiologists were significantly lower (15% in total, P < .001), with misclassification as the most prevalent category (11%). These findings align with observations by Wu et al., who noted that VLMs like GPT-4V tend to list differential diagnoses based on training data rather than performing a true diagnostic evaluation[ 21 ] . Our error analysis provided further insights into the limitations of VLMs. Errors largely fell into the following categories: (1) incorrect anatomic classification, (2) inaccurate description of imaging findings, (3) misidentification of imaging sequences, (4) hallucinated findings, and (5) overlooked pathologies. A recurrent issue was the inability of VLMs to differentiate between radiological right and left, and between different lobes; for example, pathologies in the parietal, frontal, or occipital lobes were often erroneously attributed to the temporal lobe. This observation echoes the findings of Yan et al., who reported that VLMs struggle with precise spatial localization—a critical component in medical imaging interpretation[ 15 ] . In terms of imaging sequences, most VLMs tended to default to interpreting MRI images as either T1 or T2 weighted, misidentifying susceptibility weighted images. Consequently, cases involving conditions such as amyloid angiopathy or cavernomas were consistently misinterpreted. Additionally, we observed instances of hallucinated findings, where VLMs described pathologies absent from the image—for example, a large hypodense territory in otherwise healthy brain tissue or a “dural tail sign” in a case of lymphoma. An intriguing observation was the correlation between imaging planes and specific misdiagnoses. In one case involving coronal FLAIR images, GPT-4o1-Preview erroneously suggested mesial temporal sclerosis for an image showing a subcortical Multinodular and Vacuolating Neuronal Tumor, despite normal hippocampal appearance. This study has several limitations. First, the dataset consisted of only 100 cases, relying on static image interpretation, which may not fully capture the performance of VLMs when analyzing complete 3D CT and MRI scans. Second, all models were evaluated using a single, predefined prompt and a zero-temperature setting; performance could vary with alternative prompt engineering strategies or different temperature settings[ 22 ]. Lastly, although our error classification was systematic, the inherent subjectivity in some assessments may have introduced bias. Future research should consider automated methods for evaluating AI-generated diagnoses to enhance objectivity and reproducibility. In conclusion, our findings indicate that both commercial and open-source VLMs currently lack the robust understanding of medical imaging, as well as the ability to provide correct differential diagnoses. While VLMs have made significant advancements in computer vision and natural language processing, they currently remains far from being suitable to effectively support physicians in neuroradiological practice. Declarations Author Contribution A.M. and S.S. designed the study. P.P., I.D, S.N., S.J and S.S rated the cases. A.M. and S.S. evaluared LLMs answers. A.M. did the analyses. A.M. and I.R. drafted the manuscript. All authors reviewed the final version of the Manuscript. Acknowledgement A.M. is a fellow of the BIH Charité Digital Clinician Scientist Program funded by the Charité – Universitätsmedizin Berlin and the Berlin Institute of Health at Charité (BIH). Data Availability The dataset including case presentation, images and source code for the methodology used in this study are publicly available on GitHub (https://github.com/Meddebma/VLM_Neurorad_Benchmark). References Zhang, H.; Jethani, N.; Jones, S.; Genes, N.; Major, V.J.; Jaffe, I.S.; Cardillo, A.B.; Heilenbach, N.; Ali, N.F.; Bonanni, L.J.; et al. Evaluating Large Language Models in Extracting Cognitive Exam Dates and Scores. medRxiv 2024, 2023. 07.10.23292373 , doi:10.1101/2023.07.10.23292373. Cascella, M.; Montomoli, J.; Bellini, V.; Bignami, E. Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. J. Méd. Syst. 2023, 47 , 33, doi: 10.1007/s10916-023-01925-4 . Guevara, M.; Chen, S.; Thomas, S.; Chaunzwa, T.L.; Franco, I.; Kann, B.H.; Moningi, S.; Qian, J.M.; Goldstein, M.; Harper, S.; et al. Large Language Models to Identify Social Determinants of Health in Electronic Health Records. npj Digit. Med. 2024, 7 , 6, doi: 10.1038/s41746-023-00970-0 . Meddeb, A.; Ebert, P.; Bressem, K.K.; Desser, D.; Dell’Orco, A.; Bohner, G.; Kleine, J.F.; Siebert, E.; Grauhan, N.; Brockmann, M.A.; et al. Evaluating Local Open-Source Large Language Models for Data Extraction from Unstructured Reports on Mechanical Thrombectomy in Patients with Ischemic Stroke. J. NeuroInterventional Surg. 2024, jnis-2024-022078, doi: 10.1136/jnis-2024-022078 . Veen, D.V.; Uden, C.V.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E.P.; Seehofnerová, A.; et al. Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization. Nat. Med. 2024, 30 , 1134–1142, doi: 10.1038/s41591-024-02855-5 . Meddeb, A.; Lüken, S.; Busch, F.; Adams, L.; Ugga, L.; Koltsakis, E.; Tzortzakakis, A.; Jelassi, S.; Dkhil, I.; Klontzas, M.E.; et al. Large Language Model Ability to Translate CT and MRI Free-Text Radiology Reports Into Multiple Languages. Radiology 2024, 313 , e241736, doi: 10.1148/radiol.241736 . Mohamad, F.A.; Donle, L.; Dorfner, F.; Romanescu, L.; Drechsler, K.; Wattjes, M.P.; Nawabi, J.; Makowski, M.R.; Häntze, H.; Adams, L.; et al. Open-Source Large Language Models Can Generate Labels from Radiology Reports for Training Convolutional Neural Networks. Acad. Radiol. 2025, doi: 10.1016/j.acra.2024.12.028 . Adams, L.C.; Truhn, D.; Busch, F.; Kader, A.; Niehues, S.M.; Makowski, M.R.; Bressem, K.K. Leveraging GPT-4 for Post Hoc Transformation of Free-Text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology 2023, 307 , e230725, doi: 10.1148/radiol.230725 . Schubert, M.C.; Wick, W.; Venkataramani, V. Performance of Large Language Models on a Neurology Board–Style Examination. JAMA Netw. Open 2023, 6 , e2346721, doi: 10.1001/jamanetworkopen.2023.46721 . Shu, L.; Mandel, D.; Tang, O.Y.; Jiang, Z.; Goldstein, E.; Mahta, A. Large Language Model Performance in Neurology Board Questions (S33.001). Neurology 2024, 102 , doi: 10.1212/wnl.0000000000204763 . Yang, X.; Li, T.; Wang, H.; Zhang, R.; Ni, Z.; Liu, N.; Zhai, H.; Zhao, J.; Meng, F.; Zhou, Z.; et al. Multiple Large Language Models versus Experienced Physicians in Diagnosing Challenging Cases with Gastrointestinal Symptoms. npj Digit. Med. 2025, 8 , 85, doi: 10.1038/s41746-025-01486-5 . Levra, A.G.; Gatti, M.; Mene, R.; Shiffer, D.; Costantino, G.; Solbiati, M.; Furlan, R.; Dipaola, F. A Large Language Model-Based Clinical Decision Support System for Syncope Recognition in the Emergency Department: A Framework for Clinical Workflow Integration. Eur. J. Intern. Med. 2025, 131 , 113–120, doi: 10.1016/j.ejim.2024.09.017 . Ong, J.C.L.; Jin, L.; Elangovan, K.; Lim, G.Y.S.; Lim, D.Y.Z.; Sng, G.G.R.; Ke, Y.; Tung, J.Y.M.; Zhong, R.J.; Koh, C.M.Y.; et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. arXiv 2024, doi: 10.48550/arxiv.2402.01741 . Bordes, F.; Pang, R.Y.; Ajay, A.; Li, A.C.; Bardes, A.; Petryk, S.; Mañas, O.; Lin, Z.; Mahmoud, A.; Jayaraman, B.; et al. An Introduction to Vision-Language Modeling. arXiv 2024, doi: 10.48550/arxiv.2405.17247 . Yan, Z.; Zhang, K.; Zhou, R.; He, L.; Li, X.; Sun, L. Multimodal ChatGPT for Medical Applications: An Experimental Study of GPT-4V. arXiv 2023, doi: 10.48550/arxiv.2310.19061 . Brin, D.; Sorin, V.; Barash, Y.; Konen, E.; Glicksberg, B.S.; Nadkarni, G.N.; Klang, E. Assessing GPT-4 Multimodal Performance in Radiological Image Analysis. Eur. Radiol. 2024, 1–7, doi: 10.1007/s00330-024-11035-5 . Bhayana, R.; Krishna, S.; Bleakney, R.R. Performance of ChatGPT on a Radiology Board-Style Examination: Insights into Current Strengths and Limitations. Radiology 2023, 307 , e230582, doi: 10.1148/radiol.230582 . Jin, Q.; Chen, F.; Zhou, Y.; Xu, Z.; Cheung, J.M.; Chen, R.; Summers, R.M.; Rousseau, J.F.; Ni, P.; Landsman, M.J.; et al. Hidden Flaws behind Expert-Level Accuracy of Multimodal GPT-4 Vision in Medicine. npj Digit. Med. 2024, 7 , 190, doi: 10.1038/s41746-024-01185-7 . Busch, F.; Han, T.; Makowski, M.R.; Truhn, D.; Bressem, K.K.; Adams, L. Integrating Text and Image Analysis: Exploring GPT-4V’s Capabilities in Advanced Radiological Applications Across Subspecialties. J. Méd. Internet Res. 2024, 26 , e54948, doi: 10.2196/54948 . Suh, P.S.; Shim, W.H.; Suh, C.H.; Heo, H.; Park, C.R.; Eom, H.J.; Park, K.J.; Choe, J.; Kim, P.H.; Park, H.J.; et al. Comparing Diagnostic Accuracy of Radiologists versus GPT-4V and Gemini Pro Vision Using Image Inputs from Diagnosis Please Cases. Radiology 2024, 312 , e240273, doi: 10.1148/radiol.240273 . Wu, C.; Lei, J.; Zheng, Q.; Zhao, W.; Lin, W.; Zhang, X.; Zhou, X.; Zhao, Z.; Zhang, Y.; Wang, Y.; et al. Can GPT-4V(Ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis. arXiv 2023, doi: 10.48550/arxiv.2310.09909 . Schramm, S.; Preis, S.; Metz, M.-C.; Jung, K.; Schmitz-Koep, B.; Zimmer, C.; Wiestler, B.; Hedderich, D.M.; Kim, S.H. Impact of Multimodal Prompt Elements on Diagnostic Performance of GPT-4V in Challenging Brain MRI Cases. Radiology 2025, 314 , e240689, doi: 10.1148/radiol.240689 . Additional Declarations No competing interests reported. Supplementary Files supplementaryfile.pdf Cite Share Download PDF Status: Published Journal Publication published 17 Nov, 2025 Read the published version in npj Digital Medicine → Version 1 posted Editorial decision: Revision requested 07 Jul, 2025 Reviews received at journal 01 Jul, 2025 Reviewers agreed at journal 11 Jun, 2025 Reviews received at journal 09 May, 2025 Reviewers agreed at journal 18 Apr, 2025 Reviewers invited by journal 12 Mar, 2025 Editor assigned by journal 11 Mar, 2025 Submission checks completed at journal 11 Mar, 2025 First submitted to journal 08 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6183659","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":470207726,"identity":"211b3c42-ddf8-4dcd-9ea0-cd16845840c9","order_by":0,"name":"Aymen Meddeb","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABC0lEQVRIiWNgGAWjYFCDAwlsDB8YDjAwMIN4bERqYZwB1sJMghZmHpAWBgJa+GdkJz4uYNgmx3c8+dhjm5o78ubt/AcfMJTZ4NQicSN3s/EMhtvGkmeepRvnHHtmOOcwM7MBw7k03NbcyN0mzcNwO3HDjRwz6dyGw4wzmJnZJBjbDuPUIX8jd/tvoJb6DTfyv0lbNhy2B2ph/8HY9h+nFgOgLUBf304wuJHDJs3YcDgRZAsDY9sBnFoMz7zdLM1jcNtw5pln5oY9xw4nA7UYSyScS8apRe547sbPPBW35YEh9uzBj5rDtjP4Dz788KHMDrf3BRJAzkMXTcCtARgxuB09CkbBKBgFowACAJ9bWKTK7vwPAAAAAElFTkSuQmCC","orcid":"","institution":"Hôpital Maison-Blanche, Université Reims- Champagne-Ardenne","correspondingAuthor":true,"prefix":"","firstName":"Aymen","middleName":"","lastName":"Meddeb","suffix":""},{"id":470207727,"identity":"f9f98643-8f75-4c90-b0c9-d62ae205c660","order_by":1,"name":"Ida Rangus","email":"","orcid":"","institution":"University of South Carolina","correspondingAuthor":false,"prefix":"","firstName":"Ida","middleName":"","lastName":"Rangus","suffix":""},{"id":470207728,"identity":"ffc39f34-2cb2-4336-bde7-fac436c6530d","order_by":2,"name":"Paolo Pagano","email":"","orcid":"","institution":"Hôpital Maison-Blanche, Université Reims- Champagne-Ardenne","correspondingAuthor":false,"prefix":"","firstName":"Paolo","middleName":"","lastName":"Pagano","suffix":""},{"id":470207729,"identity":"7464b36f-ca97-49da-a60d-e6cbb91572d1","order_by":3,"name":"Insaf Dekhil","email":"","orcid":"","institution":"National Institute Mongi Ben Hamida of Neurology","correspondingAuthor":false,"prefix":"","firstName":"Insaf","middleName":"","lastName":"Dekhil","suffix":""},{"id":470207730,"identity":"b0624277-f756-42d3-af2e-d9bdb9166489","order_by":4,"name":"Soumaya Jelassi","email":"","orcid":"","institution":"National Institute Mongi Ben Hamida of Neurology","correspondingAuthor":false,"prefix":"","firstName":"Soumaya","middleName":"","lastName":"Jelassi","suffix":""},{"id":470207731,"identity":"f08dc3ed-21f6-49c1-9721-27f447744887","order_by":5,"name":"Keno Bressem","email":"","orcid":"","institution":"Technical University Munich, Klinikum Rechts der Isar","correspondingAuthor":false,"prefix":"","firstName":"Keno","middleName":"","lastName":"Bressem","suffix":""},{"id":470207732,"identity":"8b8393f7-1dea-4a01-819b-4aefec1bfb86","order_by":6,"name":"Michael Scheel","email":"","orcid":"","institution":"Charité – Universitätsmedizin Berlin","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Scheel","suffix":""},{"id":470207733,"identity":"7325b454-a03e-4ec2-a0e5-468291bb0987","order_by":7,"name":"Mike P. Wattjes","email":"","orcid":"","institution":"Charité – Universitätsmedizin Berlin","correspondingAuthor":false,"prefix":"","firstName":"Mike","middleName":"P.","lastName":"Wattjes","suffix":""},{"id":470207734,"identity":"e09ce646-96d0-4612-8d46-bcb33fa3df49","order_by":8,"name":"Sonia Nagi","email":"","orcid":"","institution":"National Institute Mongi Ben Hamida of Neurology","correspondingAuthor":false,"prefix":"","firstName":"Sonia","middleName":"","lastName":"Nagi","suffix":""},{"id":470207735,"identity":"98562242-50ea-4418-9dea-5e663cc3796e","order_by":9,"name":"Laurent Pierot","email":"","orcid":"","institution":"Hôpital Maison-Blanche, Université Reims- Champagne-Ardenne","correspondingAuthor":false,"prefix":"","firstName":"Laurent","middleName":"","lastName":"Pierot","suffix":""},{"id":470207736,"identity":"a69be11f-8253-418f-adea-4a5def5dd9c8","order_by":10,"name":"Sebastien Soize","email":"","orcid":"","institution":"Hôpital Maison-Blanche, Université Reims- Champagne-Ardenne","correspondingAuthor":false,"prefix":"","firstName":"Sebastien","middleName":"","lastName":"Soize","suffix":""}],"badges":[],"createdAt":"2025-03-08 10:53:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6183659/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6183659/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41746-025-02047-6","type":"published","date":"2025-11-17T15:57:48+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":84666015,"identity":"f1177c3d-a7ea-44fd-8339-d3d426aa7a24","added_by":"auto","created_at":"2025-06-16 05:41:38","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":305906,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy comparison between neuroradiologists (mean) and various large language models (LLMs) for neuroradiological diagnosis. The purple bars represent accuracy for the single most probable diagnosis, and the yellow bars represent accuracy when considering the top three differential diagnoses.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6183659/v1/e4581b521f40ae0a16315278.jpeg"},{"id":84666021,"identity":"da100665-1c1e-4f8a-bdeb-a538054ae248","added_by":"auto","created_at":"2025-06-16 05:41:39","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":331928,"visible":true,"origin":"","legend":"\u003cp\u003eRadar chart illustrating the accuracy of five large language models (GPT-4o1-Preview, Gemini, Qwen, Llama 3.2 90b, and Grok) across different neuroradiological diagnostic categories, including infections, neoplasms, metabolic/toxic disorders, degenerative/demyelinating diseases, congenital/developmental disorders, vascular diseases, and traumatic disorders.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6183659/v1/d794ab40f8177da61dee89e8.jpeg"},{"id":84666018,"identity":"ed6fa921-09d5-410b-bff6-57f6953fa202","added_by":"auto","created_at":"2025-06-16 05:41:38","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":174169,"visible":true,"origin":"","legend":"\u003cp\u003eHarm Diagnosis Evaluation categorized by harm type. Each bar represents the percentage incidence of each harm category for evaluated models.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6183659/v1/368dacd75edf7f51c5c6714e.png"},{"id":96650132,"identity":"f546c3c4-0af4-4f62-a5e5-da7f43be93a0","added_by":"auto","created_at":"2025-11-24 16:08:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1401615,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6183659/v1/8c129db2-38d0-4904-8d7c-339cfa74dc51.pdf"},{"id":84666013,"identity":"cd2dd3cb-c671-427c-809c-8a7d06d89e25","added_by":"auto","created_at":"2025-06-16 05:41:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":275340,"visible":true,"origin":"","legend":"","description":"","filename":"supplementaryfile.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6183659/v1/bbc7ec043dc5abe082a0fb09.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"More Harm than Help? Evaluating the Capabilities of Vision-Language Models in Neurological Image Analysis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe development of Large Language Models (LLMs) has expanded the potential applications of artificial intelligence (AI) in healthcare, influencing various aspects of medical practice[\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Models such as OpenAI\u0026rsquo;s GPT, Google\u0026rsquo;s Gemini, or Meta\u0026rsquo;s Llama have demonstrated substantial potential for improving the efficiency and quality of healthcare delivery. Studies indicate that LLMs can extract information from electronic health records[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], summarize[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], translate[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] and structure medical texts[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], as well as answer medical board examination questions[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Furthermore, these models have been evaluated for their ability to analyze complex medical cases, suggest differential diagnoses[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] and support clinical decision-making processes[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] .\u003c/p\u003e \u003cp\u003eUnlike LLMs, which are exclusively trained on textual data, Vision Language Models (VLMs) leverage multimodal datasets that pair visual information, such as images or videos, with corresponding text annotations[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. The dual training enables VLMs to create joint embeddings that capture the relationships between visual and textual modalities. By combining these modalities, VLMs are applicable to tasks that require both image interpretation and contextual understanding, which are key components of radiological assessment [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] .\u003c/p\u003e \u003cp\u003eWhile initial studies have primarily assessed the diagnostic performance of VLMs in closed medical visual question-answering tasks [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], critical aspects such as the quality of generated differential diagnoses and the underlying reasoning remain underexplored[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. This study aims to evaluate the diagnostic accuracy of open-source and commercial VLMs in diagnosing a range neuroradiological diseases across various imaging modalities, comparing their performance with expert neuroradiologists. Additionally, we assess the quality of the generated differential diagnoses and perform a comprehensive error analysis, providing insights into the maturity of VLMs and their potential for clinical integration.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eThis retrospective study did use publicly available data and was therefore exempt from institutional review board approval. The inclusion of Radiopaedia cases was approved by the licensing committee under the terms of \u0026ldquo;Non-commercial use of Radiopaedia content for Machine Learning.\u0026rdquo;\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData Selection and Case Preparation\u003c/h2\u003e \u003cp\u003eA search for neuroradiological cases in the Radiopedia database (radiopedia.org) was conducted by an experienced neuroradiologist (A.M.) in December 2024, selecting cases that reflect the incidence of neuroradiological diseases in clinical practice. A total of 100 cases were included based on the following criteria: 1) adequate image quality, 2) availability of a case presentation, and 3) confirmed diagnosis. Cases were stratified by complexity into three categories: straightforward, intermediate, and challenging. For each case, between 1 and 4 representative images were chosen, and all images were downloaded in JPG format.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eLLM Inference\u003c/h3\u003e\n\u003cp\u003eThree proprietary LLMs\u0026mdash;Google\u0026rsquo;s Gemini 2.0, OpenAI\u0026rsquo;s GPT-4o1-Preview, and XAI\u0026rsquo;s Grok-2-vision \u0026mdash;and two open-source LLMs\u0026mdash;Meta\u0026rsquo;s Llama 3.2 90b and Alibaba\u0026rsquo;s Qwen\u0026mdash;were selected to analyze the cases. Models\u0026rsquo; temperature was set at 1, to ensure consistency. Access to these LLMs was obtained via API keys and Inference for all models was made between January 14th and February 4th. A standardized base prompt was used for all models and reads as follows:\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003eYou are an expert neuroradiologist. Read the case presentation: {presentation}. Analyze the images and provide the modality, a description of the images, the most probable diagnosis, and the two most important differential diagnoses.\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe source code for the methodology used in this study is publicly available on GitHub (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/Meddebma/VLM_Neurorad_Benchmark\u003c/span\u003e\u003cspan address=\"https://github.com/Meddebma/VLM_Neurorad_Benchmark\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eNeuroradiologist Rating\u003c/h3\u003e\n\u003cp\u003eFive neuroradiologists, with professional experience ranging from 8 to 17 years, were presented with the cases, including images and clinical histories, via a Google Forms questionnaire. They were instructed to provide the most probable differential diagnosis for each case.\u003c/p\u003e\n\u003ch3\u003eDiagnostic Accuracy Evaluation\u003c/h3\u003e\n\u003cp\u003eWe evaluated the accuracy of the most probable diagnoses, as well as the accuracy of all three generated diagnoses. Responses from both the LLMs and neuroradiologists were evaluated based on their accuracy in identifying the exact pathological entity. Responses were rated as incorrect if they did not match the reference diagnosis precisely (e.g., \u0026ldquo;familial cavernomatosis\u0026rdquo; was rated as incorrect in a case of \u0026ldquo;amyloid angiopathy,\u0026rdquo; despite both showing susceptibility artifacts on susceptibility-weighted imaging). However, responses providing a correct but less specific diagnosis or a more granular diagnosis than the predefined reference diagnosis were considered correct (e.g., \u0026ldquo;acute ischemic stroke\u0026rdquo; was deemed correct for a case of \u0026ldquo;ischemic stroke in the territory of the right MCA\u0026rdquo;). Furthermore, we conducted a quantitative analysis to evaluate the relationship between the number of images provided per case and the diagnostic accuracy of VLMs.\u003c/p\u003e\n\u003ch3\u003eHarmful Diagnosis Evaluation\u003c/h3\u003e\n\u003cp\u003eTo assess the potential harm of the most probable diagnoses generated by the models, we established strict evaluation criteria and performed a systematic harm assessment for each answer. A diagnosis was classified as \"harmful\" if it met at least one of the following criteria:\u003c/p\u003e \u003cp\u003e(1) Treatment delay: if a critical, time-sensitive condition such as a ruptured aneurysm, acute ischemic stroke is not identified as the correct diagnosis.\u003c/p\u003e \u003cp\u003e(2) Misclassification: If a benign lesion is misclassified as malignant, leading to undue patient anxiety and potentially unnecessary treatments, or a malignant lesion is misclassified as benign.\u003c/p\u003e \u003cp\u003e(3) Overdiagnosis: if the generated diagnosis could lead to an unnecessary invasive procedure (e.g., brain biopsy for a condition that does not require it) or an inappropriate surgical intervention, or if the model misclassifies common anatomical variants as pathological findings.\u003c/p\u003e \u003cp\u003eTwo independent raters (A.M. and S.S.) reviewed the responses and assigned harm classifications. Any discrepancies in classification were resolved through discussion between the two raters until a consensus was reached.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eError Analysis\u003c/h2\u003e \u003cp\u003eError analysis of the VLM-generated diagnoses was performed by two neuroradiologists (A.M. and S.S.), who systematically categorized the errors into five distinct types: (1) Incorrect anatomic classification (i.e., misidentification of the anatomical region or side), (2) Inaccurate description of imaging findings, (3) Misidentification of imaging sequences, (4) Hallucinated findings (i.e., interpretation of non-existent pathologies as real findings), and (5) Overlooked pathologies (failure to identify existing abnormalities Any discrepancies in classification were resolved through discussion until a consensus was reached.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eStatistical Analysis Statistical analyses were conducted in RStudio ((version 4.3.2, R Foundation). Pairwise statistical comparisons to assess diagnostic accuracy among different VLMs and neuroradiologsits was conducted using McNemar\u0026rsquo;s test. Given the multiple comparisons between neuroradiologists and the five different VLMs, Holm-Bonferroni correction was applied to adjust for multiple hypothesis testing while maintaining statistical power. The inter-rater reliability of error analysis and harmful diagnosis evaluation was evaluated using Cohen\u0026rsquo;s kappa. An adjusted P-value of \u0026lt;\u0026thinsp;.05 was considered statistically significant. For each model, we computed the Pearson correlation coefficient between the number of images and the corresponding accuracy scores using Python\u0026rsquo;s pandas library.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eData Selection\u003c/h2\u003e \u003cp\u003eA total of 100 cases were included, categorized as follows: Neoplasms (n\u0026thinsp;=\u0026thinsp;28), degenerative and demyelinating Diseases (n\u0026thinsp;=\u0026thinsp;20), vascular diseases (n\u0026thinsp;=\u0026thinsp;19), congenital and developmental disorders (n\u0026thinsp;=\u0026thinsp;15), Infections (n\u0026thinsp;=\u0026thinsp;9), traumatic disorders (n\u0026thinsp;=\u0026thinsp;5), metabolic and toxic disorders (n\u0026thinsp;=\u0026thinsp;4). Case complexities was classified as follows: straightforward (n\u0026thinsp;=\u0026thinsp;65), intermediate (n\u0026thinsp;=\u0026thinsp;25), challenging (n\u0026thinsp;=\u0026thinsp;10). Population Characteristics are presented in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePopulation Characteristics: Sex Distribution, Age (Mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD), and Number of Cases by Difficulty and Category\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003c/p\u003e \u003cp\u003e(n\u0026thinsp;=\u0026thinsp;100)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValue (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex (mean distribution)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e52 Male, 48 Female\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of pediatric cases (age\u0026thinsp;\u0026lt;\u0026thinsp;18)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAdult Age (mean\u0026thinsp;\u0026plusmn;\u0026thinsp;SD)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e46.35\u0026thinsp;\u0026plusmn;\u0026thinsp;17.36 years\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of cases (by difficulty level)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLow difficulty\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e65\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModerate difficulty\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh difficulty\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of cases (by category)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTumors\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVascular diseases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCongenital and developmental disorders\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDegenerative and demyelinating Diseases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInfections\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTraumatic disorders\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetabolic and toxic disorders\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eDiagnostic Accuracy\u003c/h2\u003e \u003cp\u003ePerformance metrics of the five VLMs (GPT-4o1-Preview, Gemini 2.0, Grok-2-vision, Qwen 2.5 and Llama 3.2 90b) compared with neuroradiologists are shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The mean diagnostic accuracy of neuroradiologists was 86.2%, significantly outperforming all LLMs. Among the evaluated models, Gemini (35%) and GPT-4o1-Preview (27%) achieved the highest accuracy, followed by Llama 3.2 90b (24%), QWEN (23%), and Grok-2-vision (9%). Pairwise statistical comparisons using McNemar\u0026rsquo;s test showed that all LLMs had significantly lower accuracy than neuroradiologists (P\u0026thinsp;\u0026lt;\u0026thinsp;.001, Holm-Bonferroni adjusted). No significant difference was observed between Gemini and GPT-4o1-Preview (P\u0026thinsp;=\u0026thinsp;.73), while Grok-2-vision was significantly worse than all other LLMs (P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Regarding the relationship between the number of images provided per case and the diagnostic accuracy, correlation analysis revealed negative relationships between the number of images and model accuracy across all models. Specifically, the Pearson correlation coefficients were as follows: GPT-4o1-Preview (-0.221), Gemini (-0.062), Qwen (-0.217), Llama 3.2 90b (-0.091), and Grok-2-vision (-0.146). Accuracies are presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and depicted in Figs.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAccuracies, harmful diagnosis categories as well as error types across five models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNeuroradiologists (mean)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGemini 2.0\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGPT-4o1-Preview\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLlama 3.2 90b\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eQwen 2.5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eGrok-2-vision\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy (most probable diagnosis)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e86,2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e35%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e27%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e23%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy (Top 3 diagnoses)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e52%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e47%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e43%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e40%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e21%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHarmful diagnoses (all categories)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e15%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e28%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e37%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e30%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e32%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e45%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTreatment Delay\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e16%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e28%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMisclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e11%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e11%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e21%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e14%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOverdiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eError Analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIncorrect anatomic classification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e29%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e51%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e43%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e53%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInaccurate description of imaging findings\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e35%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e43%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e54%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e58%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e72%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMisidentification of imaging modality/sequences\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e11%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e18%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e27%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHallucinated findings\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e19%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e23%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e42%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOverlooked pathologies\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e27%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e25%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e31%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e33%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e41%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eHarmful Diagnosis Evaluation\u003c/h2\u003e \u003cp\u003eThe clinical harm rate, defined as the proportion of generated diagnoses classified as harmful, was highest for Grok (45%), followed by GPT-4o1-Preview (37%), QWEN (32%), Llama 3.2 90b (30%), and Gemini (28%). Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e gives on overview of the percentage of misclassifications withing each category for each VLMs. Gemini 2.0 was the least harmful model with a score of 28%, with 16% in the first harm category, 11% in the second category, and 1% in the third category (Overdiagnosis). Although GPT-4o1-Preview had the second highest diagnostic accuracy, it ranked fourth in the clinical harm rate (37%), with 16% in the first category, and 21% in the second category. Grok had the highest harm rate (45%), with 28% in the first category and 17% in the second category. Holm-Bonferroni corrected pairwise comparisons showed that Grok produced significantly more harmful outputs than all other models (P\u0026thinsp;\u0026lt;\u0026thinsp;.001), while Gemini had the lowest harm rate but was not significantly different from Llama 3.2 90b (P\u0026thinsp;=\u0026thinsp;.41). In comparison, neuroradiologists had a combined harm rate of 15%, with 11% attributable to misclassification and 2% each for treatment delay and overdiagnosis. Supplementary file provides examples of cases with both correct and false diagnoses generated by the VLMs.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eError Analysis\u003c/h2\u003e \u003cp\u003eThe VLMs with varying performance levels exhibited different patterns of errors. For Gemini-2.0 and GPT-4o1-Preview, the most prevalent errors were Inaccurate description of imaging findings (35% and 43%, respectively) and Overlooked pathologies (27% and 25%, respectively). In contrast, lower-performing models such as Llama 3.2 90b, Qwen 2.5 demonstrated high Incorrect anatomic classification (51% and 43%, respectively) and high Inaccurate description of imaging findings (54% and 58%, respectively). Qwen had the highest error prevalence with the highest percentages of hallucinated findings (42%) and misidentification of imaging modality/sequences.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eClinical Vision-Language Models (VLMs) have shown promise in addressing various medical applications. However, their evaluation has primarily relied on static, structured assessments\u0026mdash;such as multiple-choice questions\u0026mdash;that fail to capture the dynamic complexities encountered in real-world clinical practice. The results of this study highlight significant shortcomings of VLMs in accurately identifying neuroradiological diagnosis based on clinical case descriptions and images. While neuroradiologists achieved a mean accuracy of 86.2%, the best-performing LLM, Gemini 2.0, only reached 35%, with other models performing even worse. Similarly, a recent study by Busch et al. found comparable results, with GPT-4V accurately identifying the primary diagnosis in 47% of cases, when cases were presented with clinical history and multiple-choice questions[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. This substantial performance gap highlights the current limitations of VLMs and suggests they are not yet reliable enough to be used as a clinical decision support system in neuroradiological practice.\u003c/p\u003e \u003cp\u003eWe assessed VLM performance at several levels. First, we examined the accuracy of the most probable diagnosis and whether the correct diagnosis was included among the top three differential diagnoses. Although all VLMs demonstrated improved accuracy when considering their top three choices, even the top-performing Gemini 2.0 only achieved 52% accuracy\u0026mdash;well below the 86.2% accuracy of neuroradiologists (P\u0026thinsp;\u0026lt;\u0026thinsp;.001). Notably, accuracy varied across disease categories, with congenital and developmental disorders yielding accuracies below 10% for all VLMs. This finding likely reflects the limited availability of high-quality, publicly accessible pediatric neuroradiology resources. Consistent with a recent study by Suh et al., our results also suggest that the unique challenges of pediatric imaging\u0026mdash;stemming from developmental variability and a diverse range of syndromic diseases\u0026mdash;are particularly problematic for these models[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] .\u003c/p\u003e \u003cp\u003eRegarding the relationship of number of images presented to the VLMs and accuracy rates, all correlations were negative, which imply that there might be a slight degradation in performance when more images are used, but the impact is modest (mostly between \u0026minus;\u0026thinsp;0.06 and \u0026minus;\u0026thinsp;0.22). This means that other factors than image quantity may play a more significant role in influencing model accuracy.\u003c/p\u003e \u003cp\u003eOur accuracy evaluation further revealed that many generated diagnoses were far from the ground truth. This misalignment raises significant concerns, as reliance on VLM-generated recommendations could lead to diagnostic errors and potential harm to patients. To quantify potential risks, we defined three harm categories: (1) treatment delay, (2) misclassification, and (3) overdiagnosis. For VLMs, treatment delay was the most frequent issue, with rates ranging from 14% for Llama 3.2 90b to 28% for Grok. Misclassification was also common, occurring in 11% of cases for Gemini 2.0 and up to 17% for GPT-4o1-Preview. Overdiagnosis was rare, with only one case each for Gemini and Qwen. In contrast, harmful diagnoses proposed by neuroradiologists were significantly lower (15% in total, P\u0026thinsp;\u0026lt;\u0026thinsp;.001), with misclassification as the most prevalent category (11%). These findings align with observations by Wu et al., who noted that VLMs like GPT-4V tend to list differential diagnoses based on training data rather than performing a true diagnostic evaluation[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] .\u003c/p\u003e \u003cp\u003eOur error analysis provided further insights into the limitations of VLMs. Errors largely fell into the following categories: (1) incorrect anatomic classification, (2) inaccurate description of imaging findings, (3) misidentification of imaging sequences, (4) hallucinated findings, and (5) overlooked pathologies. A recurrent issue was the inability of VLMs to differentiate between radiological right and left, and between different lobes; for example, pathologies in the parietal, frontal, or occipital lobes were often erroneously attributed to the temporal lobe. This observation echoes the findings of Yan et al., who reported that VLMs struggle with precise spatial localization\u0026mdash;a critical component in medical imaging interpretation[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] .\u003c/p\u003e \u003cp\u003eIn terms of imaging sequences, most VLMs tended to default to interpreting MRI images as either T1 or T2 weighted, misidentifying susceptibility weighted images. Consequently, cases involving conditions such as amyloid angiopathy or cavernomas were consistently misinterpreted. Additionally, we observed instances of hallucinated findings, where VLMs described pathologies absent from the image\u0026mdash;for example, a large hypodense territory in otherwise healthy brain tissue or a \u0026ldquo;dural tail sign\u0026rdquo; in a case of lymphoma. An intriguing observation was the correlation between imaging planes and specific misdiagnoses. In one case involving coronal FLAIR images, GPT-4o1-Preview erroneously suggested mesial temporal sclerosis for an image showing a subcortical Multinodular and Vacuolating Neuronal Tumor, despite normal hippocampal appearance.\u003c/p\u003e \u003cp\u003eThis study has several limitations. First, the dataset consisted of only 100 cases, relying on static image interpretation, which may not fully capture the performance of VLMs when analyzing complete 3D CT and MRI scans. Second, all models were evaluated using a single, predefined prompt and a zero-temperature setting; performance could vary with alternative prompt engineering strategies or different temperature settings[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Lastly, although our error classification was systematic, the inherent subjectivity in some assessments may have introduced bias. Future research should consider automated methods for evaluating AI-generated diagnoses to enhance objectivity and reproducibility.\u003c/p\u003e \u003cp\u003eIn conclusion, our findings indicate that both commercial and open-source VLMs currently lack the robust understanding of medical imaging, as well as the ability to provide correct differential diagnoses. While VLMs have made significant advancements in computer vision and natural language processing, they currently remains far from being suitable to effectively support physicians in neuroradiological practice.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eA.M. and S.S. designed the study. P.P., I.D, S.N., S.J and S.S rated the cases. A.M. and S.S. evaluared LLMs answers. A.M. did the analyses. A.M. and I.R. drafted the manuscript. All authors reviewed the final version of the Manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eA.M. is a fellow of the BIH Charit\u0026eacute; Digital Clinician Scientist Program funded by the Charit\u0026eacute; \u0026ndash; Universit\u0026auml;tsmedizin Berlin and the Berlin Institute of Health at Charit\u0026eacute; (BIH).\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe dataset including case presentation, images and source code for the methodology used in this study are publicly available on GitHub (https://github.com/Meddebma/VLM_Neurorad_Benchmark).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eZhang, H.; Jethani, N.; Jones, S.; Genes, N.; Major, V.J.; Jaffe, I.S.; Cardillo, A.B.; Heilenbach, N.; Ali, N.F.; Bonanni, L.J.; et al. Evaluating Large Language Models in Extracting Cognitive Exam Dates and Scores. \u003cem\u003emedRxiv\u003c/em\u003e 2024, 2023.\u003cdiv class=\"ExternalRefDOI\"\u003e07.10.23292373\u003c/div\u003e, doi:10.1101/2023.07.10.23292373.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCascella, M.; Montomoli, J.; Bellini, V.; Bignami, E. Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. J. M\u0026eacute;d. Syst. 2023, \u003cem\u003e47\u003c/em\u003e, 33, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10916-023-01925-4\u003c/span\u003e\u003cspan address=\"10.1007/s10916-023-01925-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuevara, M.; Chen, S.; Thomas, S.; Chaunzwa, T.L.; Franco, I.; Kann, B.H.; Moningi, S.; Qian, J.M.; Goldstein, M.; Harper, S.; et al. Large Language Models to Identify Social Determinants of Health in Electronic Health Records. npj Digit. Med. 2024, \u003cem\u003e7\u003c/em\u003e, 6, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-023-00970-0\u003c/span\u003e\u003cspan address=\"10.1038/s41746-023-00970-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeddeb, A.; Ebert, P.; Bressem, K.K.; Desser, D.; Dell\u0026rsquo;Orco, A.; Bohner, G.; Kleine, J.F.; Siebert, E.; Grauhan, N.; Brockmann, M.A.; et al. Evaluating Local Open-Source Large Language Models for Data Extraction from Unstructured Reports on Mechanical Thrombectomy in Patients with Ischemic Stroke. J. NeuroInterventional Surg. 2024, jnis-2024-022078, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jnis-2024-022078\u003c/span\u003e\u003cspan address=\"10.1136/jnis-2024-022078\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVeen, D.V.; Uden, C.V.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E.P.; Seehofnerov\u0026aacute;, A.; et al. Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization. Nat. Med. 2024, \u003cem\u003e30\u003c/em\u003e, 1134\u0026ndash;1142, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-024-02855-5\u003c/span\u003e\u003cspan address=\"10.1038/s41591-024-02855-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeddeb, A.; L\u0026uuml;ken, S.; Busch, F.; Adams, L.; Ugga, L.; Koltsakis, E.; Tzortzakakis, A.; Jelassi, S.; Dkhil, I.; Klontzas, M.E.; et al. Large Language Model Ability to Translate CT and MRI Free-Text Radiology Reports Into Multiple Languages. Radiology 2024, \u003cem\u003e313\u003c/em\u003e, e241736, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.241736\u003c/span\u003e\u003cspan address=\"10.1148/radiol.241736\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohamad, F.A.; Donle, L.; Dorfner, F.; Romanescu, L.; Drechsler, K.; Wattjes, M.P.; Nawabi, J.; Makowski, M.R.; H\u0026auml;ntze, H.; Adams, L.; et al. Open-Source Large Language Models Can Generate Labels from Radiology Reports for Training Convolutional Neural Networks. Acad. Radiol. 2025, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.acra.2024.12.028\u003c/span\u003e\u003cspan address=\"10.1016/j.acra.2024.12.028\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdams, L.C.; Truhn, D.; Busch, F.; Kader, A.; Niehues, S.M.; Makowski, M.R.; Bressem, K.K. Leveraging GPT-4 for Post Hoc Transformation of Free-Text Radiology Reports into Structured Reporting: A Multilingual Feasibility Study. Radiology 2023, \u003cem\u003e307\u003c/em\u003e, e230725, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.230725\u003c/span\u003e\u003cspan address=\"10.1148/radiol.230725\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchubert, M.C.; Wick, W.; Venkataramani, V. Performance of Large Language Models on a Neurology Board\u0026ndash;Style Examination. JAMA Netw. Open 2023, \u003cem\u003e6\u003c/em\u003e, e2346721, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2023.46721\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2023.46721\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShu, L.; Mandel, D.; Tang, O.Y.; Jiang, Z.; Goldstein, E.; Mahta, A. Large Language Model Performance in Neurology Board Questions (S33.001). Neurology 2024, \u003cem\u003e102\u003c/em\u003e, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1212/wnl.0000000000204763\u003c/span\u003e\u003cspan address=\"10.1212/wnl.0000000000204763\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, X.; Li, T.; Wang, H.; Zhang, R.; Ni, Z.; Liu, N.; Zhai, H.; Zhao, J.; Meng, F.; Zhou, Z.; et al. Multiple Large Language Models versus Experienced Physicians in Diagnosing Challenging Cases with Gastrointestinal Symptoms. npj Digit. Med. 2025, \u003cem\u003e8\u003c/em\u003e, 85, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-025-01486-5\u003c/span\u003e\u003cspan address=\"10.1038/s41746-025-01486-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevra, A.G.; Gatti, M.; Mene, R.; Shiffer, D.; Costantino, G.; Solbiati, M.; Furlan, R.; Dipaola, F. A Large Language Model-Based Clinical Decision Support System for Syncope Recognition in the Emergency Department: A Framework for Clinical Workflow Integration. Eur. J. Intern. Med. 2025, \u003cem\u003e131\u003c/em\u003e, 113\u0026ndash;120, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ejim.2024.09.017\u003c/span\u003e\u003cspan address=\"10.1016/j.ejim.2024.09.017\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOng, J.C.L.; Jin, L.; Elangovan, K.; Lim, G.Y.S.; Lim, D.Y.Z.; Sng, G.G.R.; Ke, Y.; Tung, J.Y.M.; Zhong, R.J.; Koh, C.M.Y.; et al. Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties. \u003cem\u003earXiv\u003c/em\u003e 2024, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arxiv.2402.01741\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.2402.01741\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBordes, F.; Pang, R.Y.; Ajay, A.; Li, A.C.; Bardes, A.; Petryk, S.; Ma\u0026ntilde;as, O.; Lin, Z.; Mahmoud, A.; Jayaraman, B.; et al. An Introduction to Vision-Language Modeling. \u003cem\u003earXiv\u003c/em\u003e 2024, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arxiv.2405.17247\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.2405.17247\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan, Z.; Zhang, K.; Zhou, R.; He, L.; Li, X.; Sun, L. Multimodal ChatGPT for Medical Applications: An Experimental Study of GPT-4V. \u003cem\u003earXiv\u003c/em\u003e 2023, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arxiv.2310.19061\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.2310.19061\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrin, D.; Sorin, V.; Barash, Y.; Konen, E.; Glicksberg, B.S.; Nadkarni, G.N.; Klang, E. Assessing GPT-4 Multimodal Performance in Radiological Image Analysis. Eur. Radiol. 2024, 1\u0026ndash;7, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00330-024-11035-5\u003c/span\u003e\u003cspan address=\"10.1007/s00330-024-11035-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhayana, R.; Krishna, S.; Bleakney, R.R. Performance of ChatGPT on a Radiology Board-Style Examination: Insights into Current Strengths and Limitations. Radiology 2023, \u003cem\u003e307\u003c/em\u003e, e230582, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.230582\u003c/span\u003e\u003cspan address=\"10.1148/radiol.230582\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin, Q.; Chen, F.; Zhou, Y.; Xu, Z.; Cheung, J.M.; Chen, R.; Summers, R.M.; Rousseau, J.F.; Ni, P.; Landsman, M.J.; et al. Hidden Flaws behind Expert-Level Accuracy of Multimodal GPT-4 Vision in Medicine. npj Digit. Med. 2024, \u003cem\u003e7\u003c/em\u003e, 190, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41746-024-01185-7\u003c/span\u003e\u003cspan address=\"10.1038/s41746-024-01185-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBusch, F.; Han, T.; Makowski, M.R.; Truhn, D.; Bressem, K.K.; Adams, L. Integrating Text and Image Analysis: Exploring GPT-4V\u0026rsquo;s Capabilities in Advanced Radiological Applications Across Subspecialties. J. M\u0026eacute;d. Internet Res. 2024, \u003cem\u003e26\u003c/em\u003e, e54948, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/54948\u003c/span\u003e\u003cspan address=\"10.2196/54948\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuh, P.S.; Shim, W.H.; Suh, C.H.; Heo, H.; Park, C.R.; Eom, H.J.; Park, K.J.; Choe, J.; Kim, P.H.; Park, H.J.; et al. Comparing Diagnostic Accuracy of Radiologists versus GPT-4V and Gemini Pro Vision Using Image Inputs from Diagnosis Please Cases. Radiology 2024, \u003cem\u003e312\u003c/em\u003e, e240273, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.240273\u003c/span\u003e\u003cspan address=\"10.1148/radiol.240273\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, C.; Lei, J.; Zheng, Q.; Zhao, W.; Lin, W.; Zhang, X.; Zhou, X.; Zhao, Z.; Zhang, Y.; Wang, Y.; et al. Can GPT-4V(Ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis. \u003cem\u003earXiv\u003c/em\u003e 2023, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arxiv.2310.09909\u003c/span\u003e\u003cspan address=\"10.48550/arxiv.2310.09909\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchramm, S.; Preis, S.; Metz, M.-C.; Jung, K.; Schmitz-Koep, B.; Zimmer, C.; Wiestler, B.; Hedderich, D.M.; Kim, S.H. Impact of Multimodal Prompt Elements on Diagnostic Performance of GPT-4V in Challenging Brain MRI Cases. Radiology 2025, \u003cem\u003e314\u003c/em\u003e, e240689, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/radiol.240689\u003c/span\u003e\u003cspan address=\"10.1148/radiol.240689\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6183659/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6183659/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjectives:\u003c/h2\u003e \u003cp\u003eThis study evaluates the performance of both open-source and commercial Vision Language Models (VLMs) in interpreting radiological images of neurological diseases, comparing their diagnostic accuracy to that of experienced neuroradiologists.\u003c/p\u003e\u003ch2\u003eMethods:\u003c/h2\u003e \u003cp\u003eA dataset comprising 100 cases of brain and spine pathologies with confirmed diagnoses was curated from the Radiopaedia database to reflect routine clinical neuroradiology practice. Five neuroradiologists reviewed the cases\u0026mdash;including imaging and case presentations\u0026mdash;to determine the most probable diagnosis. In parallel, five VLMs (Gemini 2.0, GPT-4o1-Preview, Llama 3.2 90b, Qwen 2.5, and Grok-2-vision) were provided with the same cases and tasked with generating three differential diagnoses along with their reasoning. Two neuroradiologists then evaluated the accuracy of both the single most probable diagnosis and the top three diagnoses produced by the VLMs, as well as the rationale provided, and assessed the potential for harmful outcomes based on the VLM outputs.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e \u003cp\u003eNeuroradiologists achieved a mean diagnostic accuracy of 86.2%, significantly outperforming all VLMs. Among the models, Gemini 2.0 achieved the highest accuracy at 35% with 28% of its diagnoses deemed potentially harmful, while Grok-2-vision had the lowest accuracy at 9% with 45% of its outputs categorized as harmful. All models demonstrated a trend toward slightly lower accuracy with an increasing number of images, however the strength of this relationship was modest. Evaluation of potential harm revealed that treatment delay was the most common risk for VLMs, ranging between 28% for Gemini 2.0 and 45% for Grok-2-vision. Error analysis indicated that the most frequent causes of misdiagnosis were incorrect anatomic classification\u0026mdash;with error rates ranging from 26% for Gemini 2.0 to 53% for Grok-2-vision \u0026mdash;and inaccurate description of imaging findings, which ranged from 35% for Gemini 2.0 to 72% for Grok-2-vision.\u003c/p\u003e\u003ch2\u003eConclusion:\u003c/h2\u003e \u003cp\u003eWhile VLMs hold promise for enhancing radiological workflows, the current state-of-the-art of open-source and commercial models is far from being reliable for the interpretation of radiological images of neurological diseases.\u003c/p\u003e","manuscriptTitle":"More Harm than Help? Evaluating the Capabilities of Vision-Language Models in Neurological Image Analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-16 05:41:33","doi":"10.21203/rs.3.rs-6183659/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-07-08T00:11:11+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-01T16:27:18+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"38136905396055581715254454791119096243","date":"2025-06-11T11:15:47+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-05-09T17:40:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"9288304214751299157110622464195469858","date":"2025-04-18T12:38:27+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-12T05:14:17+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-12T01:16:45+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-11T11:19:29+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Digital Medicine","date":"2025-03-08T10:48:45+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"npj-digital-medicine","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"npjdigitalmed","sideBox":"Learn more about [npj Digital Medicine](http://www.nature.com/npjdigitalmed/)","snPcode":"41746","submissionUrl":"https://submission.springernature.com/new-submission/41746/3","title":"npj Digital Medicine","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8d8a42d8-443e-4b11-a84f-1e1ada8da40b","owner":[],"postedDate":"June 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":49928833,"name":"Health sciences/Medical research"},{"id":49928834,"name":"Health sciences/Neurology"}],"tags":[],"updatedAt":"2025-11-24T16:01:40+00:00","versionOfRecord":{"articleIdentity":"rs-6183659","link":"https://doi.org/10.1038/s41746-025-02047-6","journal":{"identity":"npj-digital-medicine","isVorOnly":false,"title":"npj Digital Medicine"},"publishedOn":"2025-11-17 15:57:48","publishedOnDateReadable":"November 17th, 2025"},"versionCreatedAt":"2025-06-16 05:41:33","video":"","vorDoi":"10.1038/s41746-025-02047-6","vorDoiUrl":"https://doi.org/10.1038/s41746-025-02047-6","workflowStages":[]},"version":"v1","identity":"rs-6183659","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6183659","identity":"rs-6183659","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00