Benchmark Evaluations, Applications, and Challenges of Large Vision Language Models: A Survey

preprint OA: closed CC-BY-4.0
AI-generated summary by claude@2026-07, 2026-07-15

This survey systematically reviews major vision-language models, their architectures, training methods, benchmarks, applications, and challenges like hallucination and safety.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

This paper is a survey that reviews large vision-language models, focusing on benchmark evaluations, practical applications, and challenges such as performance limitations, evaluation approaches, and the tradeoffs involved in model deployment and assessment. It discusses how different benchmarks are used to measure capabilities and what kinds of evaluation difficulties researchers face, including heterogeneity across tasks and metrics. The paper does not present new original experiments, so its main limitation is that it synthesizes existing literature rather than providing direct empirical results. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Multimodal Vision Language Models (vlms) have emerged as a transformative technology at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual modalities. For example, models such as CLIP[1], Claude[2], and GPT-4V[3] demonstrate strong reasoning and understanding abilities on visual and textual data and beat classical single modality vision models on zero-shot classification[4]. Despite their rapid advancements in research and growing popularity in applications, a comprehensive survey of existing studies on vlms is notably lacking, particularly for researchers aiming to leverage vlms in their specific domains. To this end, we provide a systematic overview of vlms in the following aspects: [1] model information of the major vlms developed over the past five years (2019-2024); [2] the main architectures and training methods of these vlms; [3] summary and categorization of the popular benchmarks and evaluation metrics of vlms; [4] the applications of vlms including embodied agents, robotics, and video generation; [5] the challenges and issues faced by current vlms such as hallucination, fairness, and safety. Detailed collections including papers and model repository links are listed in https://github.com/zli12321/Awesome-VLM-Papers-And-Models.git.
Full text 621 characters · extracted from oa-doi-fallback · click to expand
There is a newer version available for this {{ publicationType }}. View latest version {{ publication.field_name }} {{ publication.subfield_name }} Copyright: © {{ publicationYear }} {{ publication.presentation_authors[0].full_name + (publication.presentation_authors.length > 1 ? ' et al' : '') }}. This is an open access publication distributed under the terms of the CC BY 4.0 License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Check the {{ publicationType | capitalize }} Source for copyright and license information. Listen on

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0