learning using privileged information

Designs training algorithms, teacher models, and distillation procedures that exploit privileged information available only at training time to produce student policies or predictors that operate without those privileged inputs; this encompasses planner-to-student distillation, privileged policy/self-distillation, privileged-to-causal conversion, behavioral cloning with DAGGER, inverse-dynamics conditioned distillation, and recurrent planner architectures. Builds and evaluates teacher–student pipelines, representation-alignment and distillation objectives, and analyzes deployment-time performance and robustness (e.g., recovery from occlusion, reduced rollout forks) when privileged context is absent.

learningusingprivilegedinformation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of effectively transferring capabilities enhanced by privileged information (PI) in multi-agent environments when such PI is inaccessible during inference and only action trajectories are observable. To tackle this, the authors propose π-Distill, a joint training framework coupled with an On-Policy Self-Distillation (OPSD) reinforcement learning strategy, enabling efficient knowledge transfer from a PI-equipped teacher model to a PI-free student model using only action trajectories as supervision. This approach departs from conventional distillation paradigms that rely on full chain-of-thought supervision, and demonstrates significant performance gains over standard supervised fine-tuning followed by reinforcement learning across multiple agent benchmarks, thereby improving the reasoning capabilities of models operating without access to privileged information.

Agentic EnvironmentsKnowledge DistillationLanguage Models

Distilling Realizable Students from Unrealizable Teachers

May 14, 2025
YK
Yujin Kim
🏛️ Cornell University

This paper addresses policy distillation under privileged information: the teacher observes the full state, whereas the student accesses only partial observations—inducing information asymmetry, distributional shift, and policy degradation. Existing approaches either degrade teacher capability to generate realizable demonstrations or compel the student to blindly explore unobserved states, both yielding low sample efficiency. We propose an active querying–correction mechanism and intelligent reinitialization to construct recoverable trajectories within the student’s observable subspace, avoiding forced imitation of unrealizable teacher policies. Our method integrates adaptive-query imitation learning with recovery-state–based reinforcement learning, requiring no teacher modification or auxiliary exploration. Evaluated on both simulation and real-robot tasks, our approach significantly improves training efficiency and final performance, consistently outperforming standard teacher–student distillation baselines.

Efficient interaction between student and teacher neededInformation asymmetry causes policy degradationStudent learns from teacher with partial observations

This work addresses a key limitation in existing on-policy self-distillation methods, which fail to effectively leverage privileged knowledge embedded in post-hoc feedback (e.g., success/failure outcomes) from student trajectories. The authors propose PAST, a novel approach that, for the first time, utilizes complete student trajectories as privileged information to adaptively refine the teacher model. While preserving the student’s distillation prefix, PAST employs trajectory-conditioned distillation to disentangle transferable policy shifts from trajectory-specific variations and theoretically characterizes the teacher’s capacity to convey knowledge to a prefix-only student. The method integrates Forward-KL distillation, student-proximity regularization, and a distribution-preserving mechanism over correct trajectories. Evaluated on three mathematical reasoning benchmarks, PAST achieves a 5.6 percentage point improvement in Avg@12 macro-average over vanilla OPSD, with ablation studies confirming the critical roles of trajectory completeness and teacher adaptivity.

on-policy self-distillationprivileged informationreasoning models

This study addresses the issue that multi-turn agent self-distillation often induces hallucinated confidence due to information loss, resulting in performance inferior to the base model. To overcome this, we propose Privileged Self-Practice (PSP), whose core innovation lies in shifting privileged information from the loss function to the sampling stage. Specifically, PSP injects privileged information into instructions and performs online policy resampling based on the GRPO objective, combined with an analyzer model to guide training, thereby effectively mitigating spurious confidence. Experimental results demonstrate that PSP comprehensively surpasses existing baselines on the AppWorld and SWE-bench benchmarks, achieving up to a 65% improvement in task completion rate.

large language modelsmulti-turn agentspost-training

This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.

On-Policy DistillationPrivileged InformationReinforcement Learning

Latest Papers

What's happening recently
View more

This study addresses the limitation of on-policy contextual distillation, where employing instance-level ground-truth answers as privileged information degrades out-of-distribution (OOD) generalization. To overcome this, we propose substituting such answers with generic formatted instructions targeting common errors as the teacher model’s privileged information, optimizing the student model via Kullback-Leibler divergence minimization. Notably, this work provides the first demonstration that concise, universal instructions exhibit superior transferability compared to information-dense, instance-specific answers. Extensive experiments across eight autoformalization tasks reveal that our approach improves OOD accuracy by 4 to 17 percentage points in seven settings while preserving in-domain performance, establishing a more effective paradigm for knowledge distillation in formal reasoning.

Knowledge DistillationOn-Policy Context DistillationOut-of-Distribution Generalization

Hot Scholars

GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl
GS

Guillaume Sartoretti

Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed Control
HX

Haoru Xue

PhD in AI Robotics, UC Berkeley
robot learningVLAhumanoid
YZ

Yuke Zhu

The University of Texas at Austin, NVIDIA Research
Robot LearningComputer VisionMachine LearningRobotics
WX

Wenli Xiao

PhD in Robotics, Carnegie Mellon University
Robot LearningReinforcement LearningHumanoids