A Policy-Graph Approach to Explain Reinforcement Learning Agents: A Novel Policy-Graph Approach with Natural Language and Counterfactual Abstractions for Explaining Reinforcement Learning Agents

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

Abstract As reinforcement learning (RL) continues to improve and be appliedin situations alongside humans, the need to explain the learned behaviorsof RL agents to end-users becomes more important. Strategies forexplaining the reasoning behind an agent’s policy, called policy-levelexplanations, can lead to important insights about both the task and theagent’s behaviors. Following this line of research, in this work, we proposea novel approach, named as CAPS, that summarizes an agent’s policy inthe form of a directed graph with natural language descriptions. A decisiontree based clustering method is utilized to abstract the state space ofthe task into fewer, condensed states which makes the policy graphs moredigestible to end-users. We then use the user-defined predicates to enrich the abstract states with semantic meaning. To introduce counterfactual state explanations to the policy graph, wefirst identify the critical states in the graph then develop a novel counterfactualexplanation method based on action perturbation in those criticalstates.We generate explanation graphs using CAPS on 5 RL tasksfor deterministic and stochastic policies. We evaluate the effectivenessof CAPS on human participants who are not RL experts in twouser studies. When provided with our explanation graph, end-users are ableto accurately interpret policies of trained RL agents 80% of the time and 68.2%of users demonstrated an increase in their confidence in understandingan agent’s behavior.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-4.0