Benchmarking Locally Deployed Open-Weight Vision–Language Models on the Japanese National Examination for Pharmacists

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Objective: To evaluate the performance of locally deployable, open-weight Vision–Language Models (VLMs) on the 2026 Japanese National Examination for Pharmacists (JNEP) and to examine the influence of model scale and generational differences on domain-specific examination outcomes. Results: All 26 open-weight VLMs were evaluated under standardized local execution conditions across three independent trials. Accuracy varied across models, with the highest-performing model achieving 62.1% (± 0.17%). Performance differed by subject domain: Pathophysiology/Drug Therapy, Biology, and Pharmacology demonstrated relatively high accuracy, whereas Chemistry tended to show lower performance than other domains across architectures and parameter scales. This pattern may reflect the difficulty of interpreting scientific imagery such as chemical structures. Model parameter count showed a moderate positive association with overall accuracy (R² = 0.59), suggesting a scaling trend in which larger models tended to achieve higher scores. Within model families, accuracy generally increased with model size. At the same time, substantial variability was observed among models, particularly below 14B. Among these, more recently released models exhibited a scale–performance relationship. These results indicate that factors beyond scale, including architectural and training differences, may contribute to performance variation. This study provides an initial benchmark of VLMs in the pharmaceutical domain.
Full text 74,098 characters · extracted from preprint-html · click to expand
Benchmarking Locally Deployed Open-Weight Vision–Language Models on the Japanese National Examination for Pharmacists | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Short Report Benchmarking Locally Deployed Open-Weight Vision–Language Models on the Japanese National Examination for Pharmacists Hiroto Asano, Yu-Shi Tian, Asuka Hatabu, Minako Ohishi, Kaori Fukuzawa, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8962575/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Objective: To evaluate the performance of locally deployable, open-weight Vision–Language Models (VLMs) on the 2026 Japanese National Examination for Pharmacists (JNEP) and to examine the influence of model scale and generational differences on domain-specific examination outcomes. Results: All 26 open-weight VLMs were evaluated under standardized local execution conditions across three independent trials. Accuracy varied across models, with the highest-performing model achieving 62.1% (± 0.17%). Performance differed by subject domain: Pathophysiology/Drug Therapy, Biology, and Pharmacology demonstrated relatively high accuracy, whereas Chemistry tended to show lower performance than other domains across architectures and parameter scales. This pattern may reflect the difficulty of interpreting scientific imagery such as chemical structures. Model parameter count showed a moderate positive association with overall accuracy (R² = 0.59), suggesting a scaling trend in which larger models tended to achieve higher scores. Within model families, accuracy generally increased with model size. At the same time, substantial variability was observed among models, particularly below 14B. Among these, more recently released models exhibited a scale–performance relationship. These results indicate that factors beyond scale, including architectural and training differences, may contribute to performance variation. This study provides an initial benchmark of VLMs in the pharmaceutical domain. large language models Vision–Language Models open-weight models pharmacist examination Japanese National Examination for Pharmacists benchmark Figures Figure 1 Figure 2 Figure 3 Introduction Recent advances in artificial intelligence, particularly in large language models (LLMs), are increasingly influencing a range of scientific and professional domains.[ 1 ] Professional qualification examinations have become one of several useful, relatively standardized benchmarks for assessing the domain-specific capabilities of LLMs. Large-scale evaluations across international licensing examinations have shown improvements in LLM performance in exam-style medical question answering.[ 2 ] In Japan, some studies suggest that LLMs can approach passing-level performance on the Japanese National Examination for Pharmacists (JNEP).[ 3 , 4 ] However, these evaluations have mainly focused on commercial, cloud-based models. While such models generally offer strong performance, concerns surrounding data privacy, service continuity, and experimental reproducibility remain important considerations, motivating interest in locally deployable, open-weight LLMs for research and clinical applications.[ 5 ] Our group has previously reported that several locally deployed LLMs can achieve accuracy rates at or near the JNEP passing threshold.[ 6 ] Alongside these findings, we have also raised concerns regarding potential data leakage—namely, the possibility that examination questions may be inadvertently included in the pretraining corpora of these models. Furthermore, prior evaluations have largely focused on text-based inputs, leaving the assessment of image-containing questions—such as those involving chemical structures and molecular diagrams—substantially underexplored. As visual representations play an important role in pharmaceutical knowledge, evaluating open-weight Vision–Language Models (VLMs) on image-containing pharmaceutical examination items is an important next step. In the present study, we conducted a cross-sectional evaluation of 26 open-weight VLMs using the JNEP administered on February 21–22, 2026. To minimize the risk of data leakage, model evaluations were performed immediately after the examination, using the original question sets and provisional answer keys from a Japanese preparatory school specializing in pharmacist examination training. All models were executed locally via the Ollama.[ 7 ] This study contributes an initial benchmark of the performance of open-weight VLMs in the pharmaceutical domain. Materials and Methods 2.1 Japanese National Examination for Pharmacists The JNEP is a national licensing examination administered annually by the Ministry of Health, Labour and Welfare of Japan. In 2026, the examination was conducted on February 21–22 (the 111th administration). The examination comprises 345 multiple-choice questions spanning nine subject domains (Table 1 ). The examination is graded on a relative basis, and the passing score varies from year to year; however, in most years, a score of approximately 60% or higher is required to pass. Table 1 Subject domains and number of questions in the 2026 Japanese National Examination for Pharmacists. Domain Number of Questions Number of Questions Including Images Physics 25 8 Chemistry 25 19 Biology 25 6 Hygiene 50 12 Pharmacology 50 1 Pharmaceutics 50 14 Pathophysiology/Drug Therapy 50 0 Laws/Regulations/Ethics 40 2 Practice a 30 2 Total 345 64 JNEP examination questions were used. As official answer keys were not yet available at the time of the study, provisional reference answers were prepared based on publicly available rapid reports from a pharmacist examination preparatory school. [ 8 ] Each question, comprising the question stem and answer choices, was individually processed; image-containing questions were extracted and stored as image files, while text-only components were stored as plain text. 2.2 Models and Computational Environment All model evaluations were performed on an NVIDIA DGX Spark using the Ollama, which enables consistent local execution and management of open-weight models. [ 7 ] A total of 26 open-weight VLMs available within the Ollama library were evaluated (Table 2 ). The qwen3 series models were not included due to their time-consuming nature. Sampling parameters were set to Ollama defaults. Table 2 open-weight VLMs evaluated in this study Region Developer Series Model Number of Parameters Release year United States Liu et al. LLaVA llava:7b 7B 2023 United States Liu et al. LLaVA llava:13b 13B 2023 SkunkworksAI BakLLaVA bakllava:7b 7B 2023 United States Moondream AI Moondream moondream:1.8b 1.8B 2024 United States Meta LLaMA Vision llama3.2-vision:11b 11B 2024 United States Meta LLaMA Vision llama3.2-vision:90b 90B 2024 United States Meta LLaMA 4 llama4:16x17b 109B 2025 United States IBM Granite Vision granite3.2-vision:2b 2B 2025 United States Google Gemma 3 gemma3:4b 4B 2025 United States Google Gemma 3 gemma3:12b 12B 2025 United States Google Gemma 3 gemma3:27b 27B 2025 United States Google TranslateGemma translategemma:4b 4B 2026 United States Google TranslateGemma translategemma:12b 12B 2026 United States Google TranslateGemma translategemma:27b 27B 2026 Europe Mistral AI Ministral ministral-3:3b 3B 2025 Europe Mistral AI Ministral ministral-3:8b 8B 2025 Europe Mistral AI Ministral ministral-3:14b 14B 2025 Europe Mistral AI Devstral devstral-small-2:24b 24B 2025 Europe Mistral AI Mistral Small mistral-small3.1:24b 24B 2025 Europe Mistral AI Mistral Small mistral-small3.2:24b 24B 2025 China XTuner LLaVA llava-llama3:8b 8B 2024 China XTuner LLaVA llava-phi3:3.8b 3.8B 2024 China Alibaba Qwen2.5-VL qwen2.5vl:3b 3B 2025 China Alibaba Qwen2.5-VL qwen2.5vl:7b 7B 2025 China Alibaba Qwen2.5-VL qwen2.5vl:32b 32B 2025 China Alibaba Qwen2.5-VL qwen2.5vl:72b 72B 2025 2.3 Prompt Design and Input Construction A standardized prompt structure was applied uniformly across all models to ensure consistency. For each item, the question stem and answer choices were first presented verbatim as in the original examination. After the full question content, explicit output instructions were added to the prompt. These instructions required the model to output only the answer number(s), to list multiple correct answers as comma-separated integers when applicable, and to refrain from generating any explanatory text or additional content. To ensure output consistency and minimize language-related variability, all instructions were written in English. 2.4 Evaluation Procedure Each model was evaluated on all examination items across three independent trials, and the mean accuracy across trials was used for subsequent analyses. Model outputs were assessed against reference answers. For items requiring multiple correct responses, ground-truth labels were normalized into an order-invariant representation by sorting the correct indices in ascending order. For questions requiring multiple correct answers, the extracted integers were treated as an order-normalized set and evaluated for exact set equality. Responses were considered correct only if the normalized answer sets matched exactly; all other outputs were recorded as incorrect. Results All 26 VLMs were evaluated on the 2026 JNEP (Fig. 1 , Supplementary Tables 1–2). Overall performance varied across models. The highest accuracy was achieved by qwen2.5-vl:72b (mean accuracy 62.1% ± 0.17% across three trials). Performance also differed substantially across subject domains. The highest mean accuracy was observed in Pathophysiology/Drug Therapy, where top-performing models reached up to 80%, followed by Biology (up to 78.7%) and Pharmacology (up to 70.7%). In contrast, Chemistry demonstrated the lowest overall performance, with a maximum accuracy of 40%, indicating that this domain posed a particular challenge across architectures and scales. A positive association was observed between parameter count and overall accuracy. Linear regression analysis showed a positive association between model size and performance (R² = 0.59; Fig. 2 ), suggesting a scaling effect in which larger models tend to achieve higher accuracy. However, this relationship was not uniform across all parameter ranges. Notably, among models with parameter counts to 14B, substantial variability in performance was observed even at matched scales. In addition, differences in scaling behavior were observed across release years. Models released in 2025 or later showed a more consistent association between scale and performance compared with earlier models. Within model families offering multiple parameter variants—including Gemma 3 (4B, 12B, 27B), TranslateGemma (4B, 12B, 27B), Ministral-3 (3B, 8B, 14B), and Qwen2.5-VL (3B, 7B, 32B, 72B)—accuracy tended to increase with model scale (Fig. 3 ). Within-family results align with the observed scaling pattern. At equivalent scales, performance differences between Gemma 3 and its translation-enhanced counterpart, TranslateGemma, were minimal at equivalent scales. Discussion In this study, we evaluated 26 VLMs on the 2026 JNEP and observed substantial variability in performance across models, subject domains, and parameter scales. Overall accuracy differed between models. This variability may reflect differences in current multimodal model capabilities and suggests the need for careful model selection when applying VLMs to domain-specific examinations spanning multiple scientific disciplines, such as national licensure tests. Performance also varied across subject domains. Pathophysiology/Drug Therapy, Biology, and Pharmacology showed comparatively higher accuracy, whereas Chemistry was consistently more challenging. These differences may be related to variation in question format and content. Pathophysiology and Pharmacology include a larger proportion of text-based, knowledge-oriented questions that involve clinical reasoning and factual integration. In contrast, Chemistry more often features structural formulas, reaction schemes, and other specialized visual representations. The relatively lower performance observed in Chemistry suggests that, even when image input is available, current VLMs may have limitations in interpreting domain-specific scientific imagery, such as chemical structures. This interpretation was further supported by an analysis of the highest-performing model (qwen2.5-vl:72b). In this model, accuracy on text-only questions was 68.1%, whereas accuracy on image-containing questions was 35.9%. The difference was statistically significant (χ²(1) = 67.35, p < 0.001), with an odds ratio of 0.26, indicating that correct responses were significantly less likely when images were present. Scientific images such as chemical structures differ substantially from the natural images commonly used in multimodal pretraining, which may partly explain this performance gap. Model scale showed a moderate positive association with performance (R² = 0.59), broadly consistent with expectations from scaling trends.[ 9 ] Larger models—particularly those with approximately 20 billion or more parameters—tended to achieve higher accuracy. However, the relationship was not uniform. Considerable variability was observed among models with similar parameter counts, especially below 14B parameters. Notably, models released in 2025 or later exhibited a clearer, more consistent scale–performance relationship, with accuracy generally increasing with parameter count. Accordingly, model release year appears to be an important contextual factor when interpreting performance differences. These findings indicate that parameter count alone does not fully account for performance variation, and that additional factors—such as architectural design and the composition of pretraining data—likely contribute. Further investigation is warranted to clarify the relative contributions of these elements. Within model families, accuracy increased with parameter count, consistent with expected scaling behavior. In a direct comparison, translation-enhanced variants such as TranslateGemma did not show substantial performance differences compared with their base Gemma counterparts on this benchmark. [ 10 , 11 ] Within the scope of this benchmark, no clear performance differences were observed between the derivative variants and their base models. Overall, the performance of VLMs on professional examinations such as the JNEP appears to be improving alongside increases in model scale and technical progress. However, persistent difficulty remains in domains requiring interpretation of specialized scientific imagery, such as chemical structures, and performance variability across subject areas was observed. Continued technical advances in VLMs and the development of domain-specific models may contribute to improved accuracy on professional examinations such as the JNEP. Limitations This study has several limitations that should be acknowledged. First, the evaluation was restricted to multiple-choice questions; free-response formats and direct assessment of clinical reasoning in real-world practice settings were not examined. Second, all inference parameters were held at framework defaults, and model-specific optimization of inference settings was not performed. Third, prompt design was standardized and was not individually optimized for each model, which may have introduced performance variability unrelated to intrinsic model capability. Fourth, while the VLMs evaluated in this study represent a practical selection of models deployable on the NVIDIA DGX Spark via the Ollama, several large-scale models were excluded due to their substantial computational requirements and practical constraints associated with local execution. Additionally, open-weight VLMs not currently available within the Ollama model library were outside the scope of this evaluation. These limitations remain to be addressed in future studies to provide a more comprehensive benchmark of open-weight VLMs in the pharmaceutical domain. Abbreviations LLMs: large language models JNEP: Japanese National Examination for Pharmacists VLMs: Vision–Language Models Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Availability of data and materials The official examination questions and provisional answer reports were accessed from Yakugaku Seminar (Igaku Academy Co., Ltd.) (https://www.yakuzemi.ac.jp/information/111_exercise/). Competing interests The authors declare that they have no competing interests. Funding This study received partial support from the OTC Self-Medication Promotion Foundation, which was awarded to Y.S.T. Authors' contributions H.A. and Y.S.T. contributed to the conception and design. H.A. evaluates VLMs. H.A. and Y.S.T. wrote the first draft of the manuscript. A.H., M.O., K.F, D.T., and K.I. confirmed the results and provided suggestions to improve the analysis. All authors have reviewed and accepted the final submitted manuscript. Acknowledgements Not applicable References Zhou H, Liu F, Gu B, Zou X, Huang J, Wu J, et al. A Survey of Large Language Models in Medicine: Progress, Application, and Challenge. 2024. https://doi.org/10.48550/arXiv.2311.05112. Zong H, Wu R, Cha J, Wang J, Wu E, Li J, et al. Large Language Models in Worldwide Medical Exams: Platform Development and Comprehensive Analysis. J Med Internet Res. 2024;26:e66114. https://doi.org/10.2196/66114. Kunitsu Y. The Potential of GPT-4 as a Support Tool for Pharmacists: Analytical Study Using the Japanese National Examination for Pharmacists. JMIR Med Educ. 2023;9:e48452. https://doi.org/10.2196/48452. Sato H, Ogasawara K, Sakurai H. Performance Evaluation of 18 Generative AI Models (ChatGPT, Gemini, Claude, and Perplexity) in 2024 Japanese Pharmacist Licensing Examination: Comparative Study. JMIR Medical Education. 2025;11:e76925. https://doi.org/10.2196/76925. Dennstädt F, Hastings J, Putora PM, Schmerder M, Cihoric N. Implementing large language models in healthcare while balancing control, collaboration, costs and security. npj Digit Med. 2025;8:143. https://doi.org/10.1038/s41746-025-01476-7. Asano H, Takaya D, Hatabu A, Ohishi M, Fukuzawa K, Ikeda K, et al. The Effectiveness of Local Fine-Tuned LLMs: Assessment of the Japanese National Examination for Pharmacists. 2025. https://doi.org/10.21203/rs.3.rs-6444534/v1. Ollama. https://ollama.com. Accessed 23 Feb 2026. Yaku-gaku Seminar. The 111th Japanese National Examination for Pharmacists: Questions, Answers, and Commentary. Yaku-gaku Seminar (Pharmacist Licensure Examination Preparatory School). 2026. https://www.yakuzemi.ac.jp/information/111_exercise/. Accessed 23 Feb 2026. Kaplan J, McCandlish S, Henighan T, Brown TB, Chess B, Child R, et al. Scaling Laws for Neural Language Models. 2020. https://doi.org/10.48550/arXiv.2001.08361. Finkelstein M, Caswell I, Domhan T, Peter J-T, Juraska J, Riley P, et al. TranslateGemma Technical Report. 2026. https://doi.org/10.48550/arXiv.2601.09012. Team G, Kamath A, Ferret J, Pathak S, Vieillard N, Merhej R, et al. Gemma 3 Technical Report. 2025. https://doi.org/10.48550/arXiv.2503.19786. Additional Declarations No competing interests reported. Supplementary Files SupplementaryTables12.xlsx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 27 Mar, 2026 Editor invited by journal 27 Feb, 2026 Editor assigned by journal 27 Feb, 2026 Submission checks completed at journal 27 Feb, 2026 First submitted to journal 24 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8962575","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Short Report","associatedPublications":[],"authors":[{"id":597291968,"identity":"338c5729-4dff-4f56-966d-df45597778a2","order_by":0,"name":"Hiroto Asano","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Hiroto","middleName":"","lastName":"Asano","suffix":""},{"id":597291976,"identity":"5e63e280-7b60-41a5-9fca-8b58a65bb51a","order_by":1,"name":"Yu-Shi Tian","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABIUlEQVRIie2RMUvDQBiG3xKIy9WuVyKJP+FKwDpU+ldyFC5LC07OASFdClkrDv6FTroGCnGpzjc4BApxcXCSDlG8E6oUL4Kb4D1w8HHw8L7fHWCx/E1cRAABnNzZXmyJflCYUtzoFwrA1CHsm2KEybgqS9QHncvZy/q0hh9MC/7cShF0ElSlURn3mS5GH+5uwosUIVuJJVVKb54jZmbFpfxN7SIn1147AV8gTrzXFK0FIGhDMapTAjmuPFKDX2WP5xuVMmxWoqMPRcd5xAVPpCh0Md6kdFdPWglJT4p+2E5pyGQljnFPR/OleZf927jqbuAPfTmq1qQe+EEmQomzwUk2nQnTix3mX7P+j88manCIMBgIkl1lh73CpFgsFsu/4x0a81qf/V7dZAAAAABJRU5ErkJggg==","orcid":"","institution":"The University of Osaka","correspondingAuthor":true,"prefix":"","firstName":"Yu-Shi","middleName":"","lastName":"Tian","suffix":""},{"id":597291977,"identity":"787da9d0-ea54-4572-a779-e0fce96381f4","order_by":2,"name":"Asuka Hatabu","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Asuka","middleName":"","lastName":"Hatabu","suffix":""},{"id":597291978,"identity":"e6c2cd67-8584-4f4d-aaac-fe8d75ba98a1","order_by":3,"name":"Minako Ohishi","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Minako","middleName":"","lastName":"Ohishi","suffix":""},{"id":597291979,"identity":"74d15e14-3fab-44c4-91dd-943d8d5970df","order_by":4,"name":"Kaori Fukuzawa","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Kaori","middleName":"","lastName":"Fukuzawa","suffix":""},{"id":597291980,"identity":"6a849825-cce5-4698-80ba-3caebd750a56","order_by":5,"name":"Daisuke Takaya","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Daisuke","middleName":"","lastName":"Takaya","suffix":""},{"id":597291981,"identity":"57b3b5f0-518b-4932-9dd1-e6563e08e0c0","order_by":6,"name":"Kenji Ikeda","email":"","orcid":"","institution":"The University of Osaka","correspondingAuthor":false,"prefix":"","firstName":"Kenji","middleName":"","lastName":"Ikeda","suffix":""}],"badges":[],"createdAt":"2026-02-25 03:23:36","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8962575/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8962575/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104399206,"identity":"f3544f35-869a-4a95-b1d9-68120a6c9680","added_by":"auto","created_at":"2026-03-11 12:05:06","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1268212,"visible":true,"origin":"","legend":"\u003cp\u003eOverall accuracy of 26 open-weight VLMs on the 2026 JNEP.\u003c/p\u003e\n\u003cp\u003eA) Overall mean accuracy (models with accuracy ≥20% are shown for readability), and B) top six models for each subject domain. Colors indicate the country/region of development (orange: China; blue: United States; green: Europe).\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-8962575/v1/56c365c23a82f2ce35924791.png"},{"id":103572593,"identity":"ce2276e0-8f94-4637-89c5-6ee33dec746c","added_by":"auto","created_at":"2026-02-27 08:34:00","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":152573,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship between model parameter count, release year, and overall JNEP accuracy.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-8962575/v1/939ab132b322837d8a53c088.png"},{"id":103572594,"identity":"aa5699a7-aca5-4f5b-8894-646aa5e761d7","added_by":"auto","created_at":"2026-02-27 08:34:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":130132,"visible":true,"origin":"","legend":"\u003cp\u003eWithin-family scaling trends across selected VLM series.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-8962575/v1/b24f0dc8af94468efc670461.png"},{"id":104407595,"identity":"db479011-461f-4973-91b7-2dcc115e9392","added_by":"auto","created_at":"2026-03-11 12:39:01","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1904552,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8962575/v1/4597b320-a408-40c0-b2f0-3b1aaf2fcfcc.pdf"},{"id":104398828,"identity":"5dda525f-a717-4978-b16d-e6279b2fb932","added_by":"auto","created_at":"2026-03-11 12:03:52","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":40099,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTables12.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8962575/v1/7f8a349bcad7b9d1c7e9ce49.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Benchmarking Locally Deployed Open-Weight Vision–Language Models on the Japanese National Examination for Pharmacists","fulltext":[{"header":"Introduction","content":"\u003cp\u003eRecent advances in artificial intelligence, particularly in large language models (LLMs), are increasingly influencing a range of scientific and professional domains.[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] Professional qualification examinations have become one of several useful, relatively standardized benchmarks for assessing the domain-specific capabilities of LLMs. Large-scale evaluations across international licensing examinations have shown improvements in LLM performance in exam-style medical question answering.[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eIn Japan, some studies suggest that LLMs can approach passing-level performance on the Japanese National Examination for Pharmacists (JNEP).[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] However, these evaluations have mainly focused on commercial, cloud-based models. While such models generally offer strong performance, concerns surrounding data privacy, service continuity, and experimental reproducibility remain important considerations, motivating interest in locally deployable, open-weight LLMs for research and clinical applications.[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eOur group has previously reported that several locally deployed LLMs can achieve accuracy rates at or near the JNEP passing threshold.[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] Alongside these findings, we have also raised concerns regarding potential data leakage\u0026mdash;namely, the possibility that examination questions may be inadvertently included in the pretraining corpora of these models. Furthermore, prior evaluations have largely focused on text-based inputs, leaving the assessment of image-containing questions\u0026mdash;such as those involving chemical structures and molecular diagrams\u0026mdash;substantially underexplored. As visual representations play an important role in pharmaceutical knowledge, evaluating open-weight Vision\u0026ndash;Language Models (VLMs) on image-containing pharmaceutical examination items is an important next step.\u003c/p\u003e \u003cp\u003eIn the present study, we conducted a cross-sectional evaluation of 26 open-weight VLMs using the JNEP administered on February 21\u0026ndash;22, 2026. To minimize the risk of data leakage, model evaluations were performed immediately after the examination, using the original question sets and provisional answer keys from a Japanese preparatory school specializing in pharmacist examination training. All models were executed locally via the Ollama.[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] This study contributes an initial benchmark of the performance of open-weight VLMs in the pharmaceutical domain.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003e \u003cb\u003e2.1 Japanese National Examination for Pharmacists\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe JNEP is a national licensing examination administered annually by the Ministry of Health, Labour and Welfare of Japan. In 2026, the examination was conducted on February 21\u0026ndash;22 (the 111th administration). The examination comprises 345 multiple-choice questions spanning nine subject domains (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The examination is graded on a relative basis, and the passing score varies from year to year; however, in most years, a score of approximately 60% or higher is required to pass.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSubject domains and number of questions in the 2026 Japanese National Examination for Pharmacists.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDomain\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of Questions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNumber of Questions Including Images\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePhysics\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemistry\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBiology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHygiene\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePharmacology\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePharmaceutics\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePathophysiology/Drug Therapy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLaws/Regulations/Ethics\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePractice\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e345\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e64\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eJNEP examination questions were used. As official answer keys were not yet available at the time of the study, provisional reference answers were prepared based on publicly available rapid reports from a pharmacist examination preparatory school. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] Each question, comprising the question stem and answer choices, was individually processed; image-containing questions were extracted and stored as image files, while text-only components were stored as plain text.\u003c/p\u003e \u003cp\u003e \u003cb\u003e2.2 Models and Computational Environment\u003c/b\u003e \u003c/p\u003e \u003cp\u003eAll model evaluations were performed on an NVIDIA DGX Spark using the Ollama, which enables consistent local execution and management of open-weight models. [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] A total of 26 open-weight VLMs available within the Ollama library were evaluated (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The qwen3 series models were not included due to their time-consuming nature. Sampling parameters were set to Ollama defaults.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eopen-weight VLMs evaluated in this study\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRegion\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDeveloper\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSeries\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNumber of Parameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRelease year\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLiu et al.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellava:7b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLiu et al.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellava:13b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSkunkworksAI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBakLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ebakllava:7b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMoondream AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMoondream\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emoondream:1.8b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.8B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeta\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaMA Vision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellama3.2-vision:11b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e11B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeta\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaMA Vision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellama3.2-vision:90b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e90B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeta\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaMA 4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellama4:16x17b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e109B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGranite Vision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003egranite3.2-vision:2b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGemma 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003egemma3:4b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGemma 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003egemma3:12b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGemma 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003egemma3:27b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e27B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTranslateGemma\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003etranslategemma:4b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2026\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTranslateGemma\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003etranslategemma:12b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2026\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnited States\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGoogle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTranslateGemma\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003etranslategemma:27b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e27B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2026\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMinistral\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eministral-3:3b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMinistral\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eministral-3:8b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMinistral\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eministral-3:14b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDevstral\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edevstral-small-2:24b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMistral Small\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emistral-small3.1:24b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEurope\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMistral AI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMistral Small\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emistral-small3.2:24b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXTuner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellava-llama3:8b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXTuner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLLaVA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ellava-phi3:3.8b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.8B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlibaba\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eQwen2.5-VL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eqwen2.5vl:3b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlibaba\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eQwen2.5-VL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eqwen2.5vl:7b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlibaba\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eQwen2.5-VL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eqwen2.5vl:32b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChina\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAlibaba\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eQwen2.5-VL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eqwen2.5vl:72b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e72B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003e2.3 Prompt Design and Input Construction\u003c/b\u003e \u003c/p\u003e \u003cp\u003eA standardized prompt structure was applied uniformly across all models to ensure consistency. For each item, the question stem and answer choices were first presented verbatim as in the original examination. After the full question content, explicit output instructions were added to the prompt. These instructions required the model to output only the answer number(s), to list multiple correct answers as comma-separated integers when applicable, and to refrain from generating any explanatory text or additional content. To ensure output consistency and minimize language-related variability, all instructions were written in English.\u003c/p\u003e \u003cp\u003e \u003cb\u003e2.4 Evaluation Procedure\u003c/b\u003e \u003c/p\u003e \u003cp\u003eEach model was evaluated on all examination items across three independent trials, and the mean accuracy across trials was used for subsequent analyses. Model outputs were assessed against reference answers. For items requiring multiple correct responses, ground-truth labels were normalized into an order-invariant representation by sorting the correct indices in ascending order. For questions requiring multiple correct answers, the extracted integers were treated as an order-normalized set and evaluated for exact set equality. Responses were considered correct only if the normalized answer sets matched exactly; all other outputs were recorded as incorrect.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eAll 26 VLMs were evaluated on the 2026 JNEP (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Supplementary Tables\u0026nbsp;1\u0026ndash;2). Overall performance varied across models. The highest accuracy was achieved by qwen2.5-vl:72b (mean accuracy 62.1% \u0026plusmn; 0.17% across three trials). Performance also differed substantially across subject domains. The highest mean accuracy was observed in Pathophysiology/Drug Therapy, where top-performing models reached up to 80%, followed by Biology (up to 78.7%) and Pharmacology (up to 70.7%). In contrast, Chemistry demonstrated the lowest overall performance, with a maximum accuracy of 40%, indicating that this domain posed a particular challenge across architectures and scales.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA positive association was observed between parameter count and overall accuracy. Linear regression analysis showed a positive association between model size and performance (R\u0026sup2; = 0.59; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), suggesting a scaling effect in which larger models tend to achieve higher accuracy. However, this relationship was not uniform across all parameter ranges. Notably, among models with parameter counts to 14B, substantial variability in performance was observed even at matched scales. In addition, differences in scaling behavior were observed across release years. Models released in 2025 or later showed a more consistent association between scale and performance compared with earlier models.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWithin model families offering multiple parameter variants\u0026mdash;including Gemma 3 (4B, 12B, 27B), TranslateGemma (4B, 12B, 27B), Ministral-3 (3B, 8B, 14B), and Qwen2.5-VL (3B, 7B, 32B, 72B)\u0026mdash;accuracy tended to increase with model scale (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Within-family results align with the observed scaling pattern. At equivalent scales, performance differences between Gemma 3 and its translation-enhanced counterpart, TranslateGemma, were minimal at equivalent scales.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we evaluated 26 VLMs on the 2026 JNEP and observed substantial variability in performance across models, subject domains, and parameter scales. Overall accuracy differed between models. This variability may reflect differences in current multimodal model capabilities and suggests the need for careful model selection when applying VLMs to domain-specific examinations spanning multiple scientific disciplines, such as national licensure tests.\u003c/p\u003e \u003cp\u003ePerformance also varied across subject domains. Pathophysiology/Drug Therapy, Biology, and Pharmacology showed comparatively higher accuracy, whereas Chemistry was consistently more challenging. These differences may be related to variation in question format and content. Pathophysiology and Pharmacology include a larger proportion of text-based, knowledge-oriented questions that involve clinical reasoning and factual integration. In contrast, Chemistry more often features structural formulas, reaction schemes, and other specialized visual representations. The relatively lower performance observed in Chemistry suggests that, even when image input is available, current VLMs may have limitations in interpreting domain-specific scientific imagery, such as chemical structures. This interpretation was further supported by an analysis of the highest-performing model (qwen2.5-vl:72b). In this model, accuracy on text-only questions was 68.1%, whereas accuracy on image-containing questions was 35.9%. The difference was statistically significant (χ\u0026sup2;(1)\u0026thinsp;=\u0026thinsp;67.35, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), with an odds ratio of 0.26, indicating that correct responses were significantly less likely when images were present. Scientific images such as chemical structures differ substantially from the natural images commonly used in multimodal pretraining, which may partly explain this performance gap.\u003c/p\u003e \u003cp\u003eModel scale showed a moderate positive association with performance (R\u0026sup2; = 0.59), broadly consistent with expectations from scaling trends.[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] Larger models\u0026mdash;particularly those with approximately 20\u0026nbsp;billion or more parameters\u0026mdash;tended to achieve higher accuracy. However, the relationship was not uniform. Considerable variability was observed among models with similar parameter counts, especially below 14B parameters. Notably, models released in 2025 or later exhibited a clearer, more consistent scale\u0026ndash;performance relationship, with accuracy generally increasing with parameter count. Accordingly, model release year appears to be an important contextual factor when interpreting performance differences. These findings indicate that parameter count alone does not fully account for performance variation, and that additional factors\u0026mdash;such as architectural design and the composition of pretraining data\u0026mdash;likely contribute. Further investigation is warranted to clarify the relative contributions of these elements.\u003c/p\u003e \u003cp\u003eWithin model families, accuracy increased with parameter count, consistent with expected scaling behavior. In a direct comparison, translation-enhanced variants such as TranslateGemma did not show substantial performance differences compared with their base Gemma counterparts on this benchmark. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] Within the scope of this benchmark, no clear performance differences were observed between the derivative variants and their base models.\u003c/p\u003e \u003cp\u003eOverall, the performance of VLMs on professional examinations such as the JNEP appears to be improving alongside increases in model scale and technical progress. However, persistent difficulty remains in domains requiring interpretation of specialized scientific imagery, such as chemical structures, and performance variability across subject areas was observed. Continued technical advances in VLMs and the development of domain-specific models may contribute to improved accuracy on professional examinations such as the JNEP.\u003c/p\u003e\n\u003ch3\u003eLimitations\u003c/h3\u003e\n\u003cp\u003eThis study has several limitations that should be acknowledged. First, the evaluation was restricted to multiple-choice questions; free-response formats and direct assessment of clinical reasoning in real-world practice settings were not examined. Second, all inference parameters were held at framework defaults, and model-specific optimization of inference settings was not performed. Third, prompt design was standardized and was not individually optimized for each model, which may have introduced performance variability unrelated to intrinsic model capability. Fourth, while the VLMs evaluated in this study represent a practical selection of models deployable on the NVIDIA DGX Spark via the Ollama, several large-scale models were excluded due to their substantial computational requirements and practical constraints associated with local execution. Additionally, open-weight VLMs not currently available within the Ollama model library were outside the scope of this evaluation. These limitations remain to be addressed in future studies to provide a more comprehensive benchmark of open-weight VLMs in the pharmaceutical domain.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003e\u003cstrong\u003e\u003cem\u003eLLMs:\u003c/em\u003e\u003c/strong\u003e large language models\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eJNEP:\u003c/em\u003e\u003c/strong\u003e Japanese National Examination for Pharmacists\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eVLMs:\u003c/em\u003e\u003c/strong\u003e Vision\u0026ndash;Language Models\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eThe official examination questions and provisional answer reports were accessed from Yakugaku Seminar (Igaku Academy Co., Ltd.) (https://www.yakuzemi.ac.jp/information/111_exercise/).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eThis study received partial support from the OTC Self-Medication Promotion Foundation, which was awarded to Y.S.T.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eH.A. and Y.S.T. contributed to the conception and design. H.A. evaluates VLMs. H.A. and Y.S.T. wrote the first draft of the manuscript. A.H., M.O., K.F, D.T., and K.I. confirmed the results and provided suggestions to improve the analysis. All authors have reviewed and accepted the final submitted manuscript.\u003c/p\u003e\n\u003ch2\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/h2\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eZhou H, Liu F, Gu B, Zou X, Huang J, Wu J, et al. A Survey of Large Language Models in Medicine: Progress, Application, and Challenge. 2024. https://doi.org/10.48550/arXiv.2311.05112.\u003c/li\u003e\n\u003cli\u003eZong H, Wu R, Cha J, Wang J, Wu E, Li J, et al. Large Language Models in Worldwide Medical Exams: Platform Development and Comprehensive Analysis. J Med Internet Res. 2024;26:e66114. https://doi.org/10.2196/66114.\u003c/li\u003e\n\u003cli\u003eKunitsu Y. The Potential of GPT-4 as a Support Tool for Pharmacists: Analytical Study Using the Japanese National Examination for Pharmacists. JMIR Med Educ. 2023;9:e48452. https://doi.org/10.2196/48452.\u003c/li\u003e\n\u003cli\u003eSato H, Ogasawara K, Sakurai H. Performance Evaluation of 18 Generative AI Models (ChatGPT, Gemini, Claude, and Perplexity) in 2024 Japanese Pharmacist Licensing Examination: Comparative Study. JMIR Medical Education. 2025;11:e76925. https://doi.org/10.2196/76925.\u003c/li\u003e\n\u003cli\u003eDennst\u0026auml;dt F, Hastings J, Putora PM, Schmerder M, Cihoric N. Implementing large language models in healthcare while balancing control, collaboration, costs and security. npj Digit Med. 2025;8:143. https://doi.org/10.1038/s41746-025-01476-7.\u003c/li\u003e\n\u003cli\u003eAsano H, Takaya D, Hatabu A, Ohishi M, Fukuzawa K, Ikeda K, et al. The Effectiveness of Local Fine-Tuned LLMs: Assessment of the Japanese National Examination for Pharmacists. 2025. https://doi.org/10.21203/rs.3.rs-6444534/v1.\u003c/li\u003e\n\u003cli\u003eOllama. https://ollama.com. Accessed 23 Feb 2026.\u003c/li\u003e\n\u003cli\u003eYaku-gaku Seminar. The 111th Japanese National Examination for Pharmacists: Questions, Answers, and Commentary. Yaku-gaku Seminar (Pharmacist Licensure Examination Preparatory School). 2026. https://www.yakuzemi.ac.jp/information/111_exercise/. Accessed 23 Feb 2026.\u003c/li\u003e\n\u003cli\u003eKaplan J, McCandlish S, Henighan T, Brown TB, Chess B, Child R, et al. Scaling Laws for Neural Language Models. 2020. https://doi.org/10.48550/arXiv.2001.08361.\u003c/li\u003e\n\u003cli\u003eFinkelstein M, Caswell I, Domhan T, Peter J-T, Juraska J, Riley P, et al. TranslateGemma Technical Report. 2026. https://doi.org/10.48550/arXiv.2601.09012.\u003c/li\u003e\n\u003cli\u003eTeam G, Kamath A, Ferret J, Pathak S, Vieillard N, Merhej R, et al. Gemma 3 Technical Report. 2025. https://doi.org/10.48550/arXiv.2503.19786.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"large language models, Vision–Language Models, open-weight models, pharmacist examination, Japanese National Examination for Pharmacists, benchmark","lastPublishedDoi":"10.21203/rs.3.rs-8962575/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8962575/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eObjective:\u003c/h2\u003e \u003cp\u003eTo evaluate the performance of locally deployable, open-weight Vision\u0026ndash;Language Models (VLMs) on the 2026 Japanese National Examination for Pharmacists (JNEP) and to examine the influence of model scale and generational differences on domain-specific examination outcomes.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e \u003cp\u003eAll 26 open-weight VLMs were evaluated under standardized local execution conditions across three independent trials. Accuracy varied across models, with the highest-performing model achieving 62.1% (\u0026plusmn;\u0026thinsp;0.17%). Performance differed by subject domain: Pathophysiology/Drug Therapy, Biology, and Pharmacology demonstrated relatively high accuracy, whereas Chemistry tended to show lower performance than other domains across architectures and parameter scales. This pattern may reflect the difficulty of interpreting scientific imagery such as chemical structures. Model parameter count showed a moderate positive association with overall accuracy (R\u0026sup2; = 0.59), suggesting a scaling trend in which larger models tended to achieve higher scores. Within model families, accuracy generally increased with model size. At the same time, substantial variability was observed among models, particularly below 14B. Among these, more recently released models exhibited a scale\u0026ndash;performance relationship. These results indicate that factors beyond scale, including architectural and training differences, may contribute to performance variation. This study provides an initial benchmark of VLMs in the pharmaceutical domain.\u003c/p\u003e","manuscriptTitle":"Benchmarking Locally Deployed Open-Weight Vision–Language Models on the Japanese National Examination for Pharmacists","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-27 08:33:55","doi":"10.21203/rs.3.rs-8962575/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-03-27T11:30:42+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-27T13:49:39+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-27T07:04:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-27T07:04:03+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Research Notes","date":"2026-02-25T03:09:23+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-research-notes","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"resn","sideBox":"Learn more about [BMC Research Notes](http://bmcresnotes.biomedcentral.com)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/resn/default.aspx","title":"BMC Research Notes","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7e8de63d-ce9c-4143-b853-ecaccc4c1924","owner":[],"postedDate":"February 27th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-27T11:38:36+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-27 08:33:55","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8962575","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8962575","identity":"rs-8962575","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-23T02:00:01.238055+00:00
License: CC-BY-4.0