Score
Designs and evaluates methods, pipelines, and policies that collect, generate, or request training data conditioned on the current learner's policy rollouts or induced contexts. This includes on‑policy augmentation procedures, conditional teacher querying, and budgeted supervision allocation intended to align training data with the learner's deployment distribution and reduce training–test context mismatch.
This study systematically investigates the feedback-to-update mechanism in on-policy distillation (OPD), where data are generated by the current policy. Framing OPD as a feedback-to-update problem, the work introduces a formula-driven categorization framework that unifies two major update pathways: distributional loss and policy gradient–style log-ratio updates. It further incorporates novel perspectives from temporal credit assignment and temporal vocabulary routing. By leveraging techniques such as KL divergence orientation, generalized advantage estimation (GAE), and counterfactual routing, the analysis reveals that OPD performance critically depends on state compatibility and support set construction. The paper establishes a comprehensive analytical framework for OPD, derives explicit bias bounds, and proposes new methods—GAE-OPD and CR-OPD—to enhance training stability, alongside actionable diagnostic tools and a practical implementation checklist.
This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.
This work addresses the distribution mismatch commonly faced by large language model agents in supervised fine-tuning, where training relies on complete teacher demonstrations while testing depends on student-generated contexts. The authors formulate online policy data construction as a budget allocation problem and propose replacing lengthy or costly filtered teacher trajectories with a small number of unfiltered, short-step teacher continuations, strategically injected into student-induced critical contexts. By systematically exploring the design space of rollout policies, switching time distributions, continuation lengths, and filtering rules—and incorporating a dual-cost model accounting for both teacher inference and supervision signal retention—the method demonstrates strong empirical performance on HotpotQA, ALFWorld, and Terminal-Bench-Dev. Notably, it matches or exceeds existing critical-context filtering baselines on the first two benchmarks at lower computational cost, indicating that a few well-placed teacher steps can substantially enhance training efficiency.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
Standard online policy distillation (OPD) often suffers from training instability due to high noise in teacher trajectories and large variance in supervision signals. To address this, this work proposes the BRTS framework, which introduces a novel multi-trajectory sampling and prioritization mechanism that selects high-quality teacher trajectories primarily based on their correctness and secondarily on their alignment with the student’s behavior. Additionally, BRTS incorporates a ground-truth conditioned recovery strategy to handle challenging samples and integrates an auxiliary supervision loss to enhance training stability. Evaluated on demanding mathematical reasoning benchmarks—including AIME 2024/2025 and AMC 2023—BRTS significantly outperforms standard OPD, with the most pronounced gains observed on the hardest problem subsets.
Current data generation heavily relies on manual analysis of model weaknesses and hand-crafted training examples; even with LLM-based annotation, human intervention remains essential for interpreting student feedback and curating data. Method: We propose DataEnvGym—the first closed-loop teacher environment testbed designed specifically for data-generation agents—formulating data creation as a sequential decision-making task guided by student feedback. Contribution/Results: (1) A novel three-layer structured “teacher environment” framework enabling decoupled state and action spaces; (2) Integration of skill-driven, interpretable curriculum control with cross-domain generalization evaluation (math, code, VQA, tool use); (3) End-to-end integration of LLM-based data generation policies, iterative student training–evaluation–feedback loops, and multi-granularity skill representations. Experiments demonstrate sustained cross-task student performance improvement and reveal the critical impact of environmental structure on skill-teaching depth—establishing a reproducible benchmark for data-generation agent research.
This work investigates the pre-warming phase in on-policy distillation (OPD), which significantly impacts performance yet lacks a clear mechanistic understanding. The study reveals that the core of effective pre-warming lies in transferring inference patterns compatible with the teacher model, rather than merely relying on ground-truth labels. To this end, the authors propose Simple-OPD, a method that leverages teacher-generated chain-of-thought data to initialize student training via low-rank adaptation (LoRA) in a plug-in manner, followed by standard OPD training. Extensive experiments demonstrate that Simple-OPD consistently outperforms full-parameter supervised fine-tuning across diverse settings, exhibiting both strong effectiveness and robustness.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.