Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of knowledge distillation in long-horizon agent tasks caused by uniform token-level supervision. To this end, we propose LENS-OPD, a novel framework that introduces a pioneering nested coarse-to-fine supervision allocation mechanism. Through hierarchical strategies encompassing localization, verification, and refinement, this approach concentrates supervision on the intersection of future utility and learnability, dynamically adapting to the evolving capabilities of the student model. Furthermore, the framework integrates online policy distillation, trajectory exposure adjustment, and conflict-focused token-level supervision to effectively optimize the knowledge transfer process. Experimental results demonstrate that LENS-OPD significantly outperforms existing baseline methods across multiple benchmarks, yielding substantial improvements in task performance.
📝 Abstract
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Long-horizon agentic tasks
Supervision allocation
Large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Hierarchical supervision allocation
Long-horizon agents
Coarse-to-fine framework
Large language models
Y
Yuhao Sun
Ant Group
B
Binrui Wu
Alibaba International Digital Commerce Group
Zhuoer Xu
Zhuoer Xu
Nanjing University
Adversarial LearningNeural Architecture SearchFeature Engineering
M
Ming Wen
Ant Group
H
Haoxiang Xu
University of Science and Technology of China
Bin Chen
Bin Chen
Peking University
Image Super-ResolutionImage Compressive Sensing
Y
Yan Lin
University of Electronic Science and Technology of China
Q
Qianzijing Zhang
University of Science and Technology of China