Explainable Automatic Evaluation: Interpretable Decision-Making in Large Language Model Autoraters

preprint OA: closed
View at publisher

Abstract

The deployment of Large Language Models (LLMs) as automatic evaluators—termed "LLM autoraters"—has emerged as a scalable alternative to human judgment for assessing natural language generation quality. However, the opacity of LLM decision-making processes fundamentally undermines trust, debuggability, and regulatory compliance in high-stakes evaluation contexts. This study introduces a framework for explainable automatic evaluation that renders autorater decision-making transparent, interpretable, and auditable without sacrificing evaluation accuracy. Building on foundational autorater architectures and their role in Reinforcement Learning from Human Feedback, the proposed framework integrates three complementary explainability mechanisms: attention-based rationale generation that highlights input segments driving evaluation scores, counterfactual contrastive explanations that reveal how minimal input perturbations would alter evaluations, and natural language justification generation with confidence calibration. The framework operates across multiple evaluation paradigms—reference-based, reference-free, and multi-turn dialogue assessment—and supports both pairwise and pointwise scoring. Empirical evaluation on four benchmark datasets (SummEval, Topical-Chat, MT-Bench, and a newly contributed multi-aspect quality dataset) demonstrates that explainable autoraters achieve high correlation with human judgments while producing explanations that are rated highly for plausibility and faithfulness by human annotators. Crucially, the framework enables detection of systematic evaluation biases, including length bias and positional bias in multi-turn contexts, facilitating autorater calibration and debugging. The explainability mechanisms add modest computational overhead compared to black-box autoraters but enable error diagnosis that would otherwise require thousands of human annotations. The study contributes a taxonomy of explainability requirements for LLM evaluation tasks, a modular framework for generating faithful, plausible, and actionable explanations, quantitative metrics for explanation quality assessment, and open-source implementation with demonstrated bias detection and correction workflows.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00