Evaluating the strengths and weaknesses of large language models in answering neurophysiology questions
preprint
OA: closed
Abstract
Abstract Background: Large language models (LLMs), such as ChatGPT, Google's Bard, and Anthropic's Claude, demonstrate impressive natural language capabilities. Assessing their competence in specialized domains such as neurophysiology is important for determining their utility in research, education, and clinical applications. Objectives:This study evaluates and compares the performance of LLMs in answering neurophysiology questions in English and Persian across different topics and cognitive levels. Methods:Twenty questions spanning 4 topics (general, sensory system, motor system, and integrative) and 2 cognitive levels (lower-order and higher-order) were presented to the LLMs. Physiologists scored the essay-style responses from 0-5 points. Statistical analysis compared the scores at themodel, language, topic, and cognitive levels. Results:Overall,the models performed well (mean score=3.56/5), with no significant difference between language or cognitive levels. Performance was the strongest in themotor system (mean=4.52) and the weakest in integrative topics (mean=2.1). Detailed qualitative analysis revealed inconsistencies and gaps in reasoning. Conclusions: Thisstudy provides insights into LLMs’ capabilities and limitations in neurophysiology. The models exhibit competence in fundamental concepts but face challenges in advanced reasoning and integration. Targeted training could address gaps in knowledge and causal reasoning. As LLMs evolve, rigorous domain-specific assessments will be important to gauge progress.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00