on-policy data augmentation

Designs and evaluates methods, pipelines, and policies that collect, generate, or request training data conditioned on the current learner's policy rollouts or induced contexts. This includes on‑policy augmentation procedures, conditional teacher querying, and budgeted supervision allocation intended to align training data with the learner's deployment distribution and reduce training–test context mismatch.

on-policydataaugmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.

Exposure BiasImitation LearningKnowledge Distillation

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the distribution mismatch commonly faced by large language model agents in supervised fine-tuning, where training relies on complete teacher demonstrations while testing depends on student-generated contexts. The authors formulate online policy data construction as a budget allocation problem and propose replacing lengthy or costly filtered teacher trajectories with a small number of unfiltered, short-step teacher continuations, strategically injected into student-induced critical contexts. By systematically exploring the design space of rollout policies, switching time distributions, continuation lengths, and filtering rules—and incorporating a dual-cost model accounting for both teacher inference and supervision signal retention—the method demonstrates strong empirical performance on HotpotQA, ALFWorld, and Terminal-Bench-Dev. Notably, it matches or exceeds existing critical-context filtering baselines on the first two benchmarks at lower computational cost, indicating that a few well-placed teacher steps can substantially enhance training efficiency.

cost-efficient supervisiondistribution mismatchon-policy data augmentation

This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.

auditable trainingcontrol planehuman-in-the-loop

Standard online policy distillation (OPD) often suffers from training instability due to high noise in teacher trajectories and large variance in supervision signals. To address this, this work proposes the BRTS framework, which introduces a novel multi-trajectory sampling and prioritization mechanism that selects high-quality teacher trajectories primarily based on their correctness and secondarily on their alignment with the student’s behavior. Additionally, BRTS incorporates a ground-truth conditioned recovery strategy to handle challenging samples and integrates an auxiliary supervision loss to enhance training stability. Evaluated on demanding mathematical reasoning benchmarks—including AIME 2024/2025 and AMC 2023—BRTS significantly outperforms standard OPD, with the most pronounced gains observed on the hardest problem subsets.

High-VarianceOn-Policy DistillationReasoning

DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback

Oct 08, 2024
ZK
Zaid Khan
🏛️ University of North Carolina at Chapel Hill

Current data generation heavily relies on manual analysis of model weaknesses and hand-crafted training examples; even with LLM-based annotation, human intervention remains essential for interpreting student feedback and curating data. Method: We propose DataEnvGym—the first closed-loop teacher environment testbed designed specifically for data-generation agents—formulating data creation as a sequential decision-making task guided by student feedback. Contribution/Results: (1) A novel three-layer structured “teacher environment” framework enabling decoupled state and action spaces; (2) Integration of skill-driven, interpretable curriculum control with cross-domain generalization evaluation (math, code, VQA, tool use); (3) End-to-end integration of LLM-based data generation policies, iterative student training–evaluation–feedback loops, and multi-granularity skill representations. Experiments demonstrate sustained cross-task student performance improvement and reveal the critical impact of environmental structure on skill-teaching depth—establishing a reproducible benchmark for data-generation agent research.

Automating data generation for model training using autonomous agents.Creating environments for iterative feedback-driven data generation.Improving student model performance through structured teacher environments.

Latest Papers

What's happening recently
View more

This work investigates the pre-warming phase in on-policy distillation (OPD), which significantly impacts performance yet lacks a clear mechanistic understanding. The study reveals that the core of effective pre-warming lies in transferring inference patterns compatible with the teacher model, rather than merely relying on ground-truth labels. To this end, the authors propose Simple-OPD, a method that leverages teacher-generated chain-of-thought data to initialize student training via low-rank adaptation (LoRA) in a plug-in manner, followed by standard OPD training. Extensive experiments demonstrate that Simple-OPD consistently outperforms full-parameter supervised fine-tuning across diverse settings, exhibiting both strong effectiveness and robustness.

chain-of-thoughtOn-policy Distillationstudent initialization

This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.

AI governancefoundation modelsmodel modification

Hot Scholars

YL

Yaojie Lu

Institute of Software, Chinese Academy of Sciences
Information ExtractionLarge Language Models
LS

Le Sun

Institute of Software, CAS
information_retrievalnatural_language_processing
HL

Hongyu Lin

Institute of Software, Chinese Academy of Sciences
Natural Language ProcessingInformation Extraction and Machine Learning
WL

Wenjie Li

The Hong Kong Polytechnic University
Text SummarizationNatural Language UnderstandingNatural Language Generation
BH

Bingxiang He

Second year PhD Candidate, Tsinghua University
Natural Language Processing