Structured Modeling and Representation Methods for Post-Retrieval Inference Processes in Large Video Language Models

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Existing Video-RAG systems often concatenate retrieved segments directly into input, leading toreasoning drift when hard negative samples are introduced. This paper proposes a Structured Post-Retrieval Reasoning (SPRR) module for Large Video Language Models (LVLMs), explicitly modelingthe post-retrieval process into three stages:(1) Evidence Validation: Generates "decidable" sub-problems (3–8) for Top-k=20 candidate clips, outputs binary/numeric scores, and filters to k′=4–6;(2) Conflict Resolution: Establishes consistency constraints (e.g., temporal order, entity attributeinvariance) for contradictory information across multiple clips, selecting the minimum conflictsubset to form a coherent evidence pool;(3) Temporal Aggregation: Indexed by event timestamps,evidence is serialized to generate interpretable reasoning chains (including referenced clip IDs andtemporal ranges).Evaluated on MLVU (3,102 QA) and LongVideoBench (6,678 MCQ) using open-ended and multiple-choice formats respectively, while measuring interpretability metrics (averageevidence count, conflict rate, reasoning chain length) and efficiency metrics (input tokens/reasoningsteps). This validates SPRR's benefits in "reducing noise, enhancing interpretability, and improvingstability.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0