Towards Explainable RAG: Interpreting the Influence of Retrieved Passages on Generation
preprint
OA: closed
CC-BY-4.0
Abstract
Augmented Generation (RAG) models have demonstrated substantial improvements in natural language processing by incorporating external knowledge via document retrieval. However, their interpretability remains limited. Users cannot easily understand how specific retrieved documents contribute to the generated response. In this paper, we propose a framework that improves transparency in RAG by analyzing and interpreting the influence of retrieved documents. We focus on three key components: user embedding profiling, custom reward function design, and a soft prompt compression mechanism. Through comprehensive experiments using benchmark datasets, we introduce new evaluation metrics to assess source attribution and influence alignment. Our findings suggest that interpretability can be meaningfully improved without sacrificing generation quality.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-06-05T02:00:03.366016+00:00
License: CC-BY-4.0