Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the supervision mismatch problem in long-horizon agents caused by the temporal misalignment between privileged feedback and decision timestamps. To this end, we propose AlignOPSD, a novel framework that introduces a "align-then-attribute" paradigm. Specifically, it achieves context calibration across sibling rollbacks through decision-alignment supervision correction, and optimizes hierarchical credit assignment by modeling variable-duration decision spans via semi-Markov processes. By integrating online self-distillation with Qwen large language models for reinforcement learning, our method comprehensively outperforms GRPO baselines on benchmarks such as ALFWorld, yielding performance improvements of 5.5%–8.7%.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Problem

Research questions and friction points this paper is trying to address.

Long-Horizon Agents
Reinforcement Learning
On-Policy Self-Distillation
Decision-Timestamp Mismatch
Credit Assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision-Aligned Supervision
Semi-Markov Hierarchical Credit Assignment
On-Policy Self-Distillation
Long-Horizon Agents
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mingju Chen
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; School of Artificial Intelligence, Beihang University
C
Can Lv
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; School of Artificial Intelligence, Beihang University
J
Jinrong Liu
School of Artificial Intelligence, Beihang University
H
Huan Zhang
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; School of Artificial Intelligence, Beihang University
Heng Chang
Heng Chang
Tsinghua University
Trustworthy AIGraph Representation LearningData Mining
Shiji Zhou
Shiji Zhou
Associate Professor, Beihang University
Online LearningStochastic OptimizationMulti-Objective OptimizationMulti-task Learning