TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

📅 2026-08-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge arising from sparse terminal rewards in long-horizon LLM agents by proposing TRCA. This method introduces a novel anchor-free, fine-grained credit assignment mechanism that eliminates reliance on process reward models or successful trajectories. Instead, it generates step-level advantage signals directly from state transition evidence, execution outcomes, and invalidity criteria to optimize policies. Experiments demonstrate that TRCA significantly outperforms baselines on benchmarks such as ALFWorld, achieving score improvements of 6.0%–12.6% on WebShop and average increases of 1.9%–18.3% on SearchQA. These results confirm its effectiveness in enabling precise policy optimization for long-horizon tasks without requiring dense supervision or expert demonstrations.
📝 Abstract
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from successful trajectories. However, successful trajectories are extremely scarce during early-stage reinforcement learning, substantially weakening anchor-based methods. We propose Transition-wise Rubric Credit Assignment (TRCA), which derives step-level supervision directly from action-induced transitions without learned evaluators or successful anchors. TRCA evaluates each transition using Evidence, Execution, and Invalidity rubrics to capture task-relevant information acquisition, valid task execution, and invalid or regressive behavior. From these judgments, Foundational Rubric Reward measures local transition quality, while Breakthrough Rubric Reward tracks newly covered Evidence and Execution conditions to reward incremental task progress. Combined with terminal outcomes, these signals produce fine-grained step-level advantages for policy optimization. Experiments on ALFWorld, WebShop, and seven search-augmented question-answering benchmarks show consistent improvements over the evaluated baselines. With Qwen2.5-7B-Instruct, TRCA improves the WebShop score by 6.0%-12.6%; with Qwen2.5-3B-Instruct, it improves the average SearchQA score by 1.9%-18.3%. These results demonstrate the effectiveness of transition-wise rubric credit assignment for long-horizon tasks with sparse successful anchors.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon LLM Agents
Credit Assignment
Sparse Rewards
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Credit Assignment
Transition-wise Rubric
Long-horizon LLM Agents
Step-level Supervision
Sparse Reward
🔎 Similar Papers
No similar papers found.
H
Huan Zhang
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
M
Mingju Chen
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
D
Dongxu Zhou
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
C
Can Lv
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
Heng Chang
Heng Chang
Tsinghua University
Trustworthy AIGraph Representation LearningData Mining
Sen Cui
Sen Cui
Tsinghua Universitty
trust LLMAI Agentembodied intelligence
F
Faguo Wu
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
Shiji Zhou
Shiji Zhou
Associate Professor, Beihang University
Online LearningStochastic OptimizationMulti-Objective OptimizationMulti-task Learning