SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse rewards obscuring intermediate contributions and distillation discarding teacher guidance in long-horizon agent training. To this end, we propose a framework that decouples planning from execution. Methodologically, long-horizon interactions are decomposed into subtasks, and a cross-rollout prefix tree is introduced to enable structured credit assignment for optimizing planning. Concurrently, local context distillation is employed to preserve teacher guidance signals, thereby enhancing execution. By integrating outcome-oriented credit assignment with teacher-guided distillation, this framework effectively disentangles and synergistically improves both planning and execution capabilities. Experimental results demonstrate that the proposed approach yields macro-average accuracy improvements of 4.48 and 4.19 percentage points on text-based and multimodal benchmarks, respectively.
📝 Abstract
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
Problem

Research questions and friction points this paper is trying to address.

Long-Horizon Agents
Credit Assignment
Sparse Rewards
Knowledge Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Credit Assignment
Knowledge Distillation
Long-Horizon Agents
Subtask Prefix Trees
On-policy Distillation
S
Shangyang Wu
Beijing University of Posts and Telecommunications
S
Shuai Zhao
Beijing University of Posts and Telecommunications
Z
Ziyue Zhu
Beijing University of Posts and Telecommunications
J
Jinyang Wu
Tsinghua University
A
Anh Tuan Luu
Nanyang Technological University, Singapore
Haoran Luo
Haoran Luo
Nanyang Technological University
Knowledge GraphLarge Language ModelsGraph Neural Networks