Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost and low efficiency of evaluating large language model (LLM) agent trajectories in production environments by proposing LiteTrajEval, a novel lightweight, budget-constrained trajectory evaluation paradigm. The method integrates offline rule-based profiling with online preprocessing, combining heuristic failure signal detection and global sequential compression techniques. Under a fixed computational budget, it leverages a single scoring-guideline LLM judge to generate structured diagnostic reports. Experimental results demonstrate that this approach improves fault localization alignment by 20–35 percentage points while reducing evaluation costs by approximately sixfold and decreasing assessment latency by over eightfold. Furthermore, LiteTrajEval has been successfully deployed within an enterprise platform, validating its practical effectiveness for scalable and efficient trajectory evaluation in real-world applications.
📝 Abstract
Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textit{LiteTrajEval}, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and $τ$-bench-style trajectory datasets, LiteTrajEval improves failure-localization alignment with human annotations by roughly 20--35 percentage points on Magentic-One and up to 23 percentage points on $τ$-retail compared with AgentRx, while reducing cost by about 6$\times$ and evaluation time by more than 8$\times$. This solution has also been deployed in our enterprise agentic platform.
Problem

Research questions and friction points this paper is trying to address.

trajectory evaluation
LLM-based agents
production environment
cost reduction
failure localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory Evaluation
Lightweight Architecture
Budget-Bounded
Rubric-Guided LLM Judge
Failure Localization
🔎 Similar Papers
No similar papers found.