Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems

๐Ÿ“… 2026-02-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the limitations of existing RAG evaluation frameworks in identifying critical failure modes in enterprise-grade multi-turn dialoguesโ€”such as case misidentification, workflow misalignment, and partial resolution across turns. To this end, we propose a case-aware evaluation framework that, for the first time, incorporates case-workflow alignment as a core evaluation dimension. The framework introduces eight operation-oriented metrics for fine-grained analysis of each conversational turn and employs a severity-aware scoring mechanism to mitigate score inflation and enhance diagnostic precision. Built upon an LLM-as-a-Judge architecture with deterministic prompting and strict JSON output formatting, our approach enables interpretable, production-ready, and batch-deployable evaluation. Experimental results demonstrate that the framework effectively uncovers key performance trade-offs in enterprise settings that are invisible to conventional agent-level metrics, thereby delivering actionable insights for system optimization.

Technology Category

Multiagent Systems: Agreement, Argumentation & NegotiationMachine Learning: Evaluation and AnalysisNatural Language Processing: Conversational AI/Dialog Systems

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
๐Ÿ“ Abstract
Enterprise Retrieval-Augmented Generation (RAG) assistants operate in multi-turn, case-based workflows such as technical support and IT operations, where evaluation must reflect operational constraints, structured identifiers (e.g., error codes, versions), and resolution workflows. Existing RAG evaluation frameworks are primarily designed for benchmark-style or single-turn settings and often fail to capture enterprise-specific failure modes such as case misidentification, workflow misalignment, and partial resolution across turns. We present a case-aware LLM-as-a-Judge evaluation framework for enterprise multi-turn RAG systems. The framework evaluates each turn using eight operationally grounded metrics that separate retrieval quality, grounding fidelity, answer utility, precision integrity, and case/workflow alignment. A severity-aware scoring protocol reduces score inflation and improves diagnostic clarity across heterogeneous enterprise cases. The system uses deterministic prompting with strict JSON outputs, enabling scalable batch evaluation, regression testing, and production monitoring. Through a comparative study of two instruction-tuned models across short and long workflows, we show that generic proxy metrics provide ambiguous signals, while the proposed framework exposes enterprise-critical tradeoffs that are actionable for system improvement.
Problem

Research questions and friction points this paper is trying to address.

RAG evaluation
enterprise RAG
multi-turn dialogue
case-aware evaluation
workflow alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

case-aware evaluation
LLM-as-a-Judge
enterprise RAG
multi-turn workflows
severity-aware scoring
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mukul Chhabra
Dell Technologies
L
Luigi Medrano
Dell Technologies
A
Arush Verma
Dell Technologies