VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos

preprint OA: closed
Full text JSON View at publisher

Abstract

Synthetic video generation using foundation models has gained significant attention due to its realism and broad applications. However, while these models excel at generating visually coherent and high-quality video frames, they often overlook commonsense reasoning and physical law violations, leading to abnormal content. Existing score-based evaluations like VideoScore[1] mainly focus on general video quality and do not take these abnormalities into account, and offer no explanations of the evaluation results. A more promising evaluation approach is to leverage multi-modal large language models (MLLMs) as interpretable video evaluators, following the approach of FactScore[2]. However, how well MLLMs can detect these abnormalities in synthetic videos is underexplored. Motivated by a more interpretable video generation evaluation, we introduce VideoHallu, a benchmark built from synthetic videos produced by popular models like Veo2[3], Sora[4], and Kling[5], paired with expert-crafted question-answering pair examples easily solvable with human-level perception and reasoning across multiple categories. We evaluate several State-of-the-Art (SoTA) MLLMs with our benchmark, including GPT-4o[6], Gemini-2.5-Pro[7], Qwen-2.5-VL[8], and forefront models like Video-R1[9] and VideoChat-R1[10]. Despite the strong performance of R1 MLLMs on real-world video benchmarks like MVBench[11] and MovieChat[12], these models still struggle and hallucinate on basic commonsense and physics reasoning tasks in synthetic videos, highlighting synthetic video hallucination as an underexplored challenge. Moreover, we post-train current SoTA MLLMs with Group Relative Policy Optimization (GRPO)[13] using both real-world and synthetic commonsense/physics datasets. Our results show improved overall accuracy compared to the base model, achieving the highest performance among all models, highlighting the importance of integrating high-quality counterexamples to enhance commonsense and physics reasoning in MLLMs’ language priors. Our data is available at https://github.com/zli12321/VideoHallu.git.
Full text 621 characters · extracted from oa-doi-fallback · click to expand
There is a newer version available for this {{ publicationType }}. View latest version {{ publication.field_name }} {{ publication.subfield_name }} Copyright: © {{ publicationYear }} {{ publication.presentation_authors[0].full_name + (publication.presentation_authors.length > 1 ? ' et al' : '') }}. This is an open access publication distributed under the terms of the CC BY 4.0 License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Check the {{ publicationType | capitalize }} Source for copyright and license information. Listen on

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00