hybrid teacher distillation

Designs and implements distillation pipelines that combine supervisory signals from one or more teacher models—including prompt-based, segmentation, BEV, policy, local/LBS, and multi-source teachers—into a student model via losses, prompts, and steering mechanisms (e.g., steered-teacher, local-teacher, or unsupervised prompt strategies). Builds methods to extract and transfer soft structural priors and hybrid supervision (fitted plus unlabeled data), and analyzes training protocols, end-to-end behavior, and scalability when training students without explicit fitting or on limited fitted datasets.

hybridteacherdistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Who Taught You That? Tracing Teachers in Model Distillation

Feb 10, 2025
SW
Somin Wadhwa
🏛️ Northeastern University

This paper addresses the “teacher attribution” problem in model distillation: Can the teacher large language model (LLM) used for distillation be identified solely from the student model’s outputs? To tackle this, we propose a novel pedagogical fingerprint based on part-of-speech (PoS) templates—revealing that PoS sequences in student outputs robustly inherit teacher-specific syntactic preferences and exhibit superior discriminability and robustness over conventional n-gram similarity. Under a black-box teacher assumption, we design a lightweight discriminative model that jointly encodes PoS templates and lexical statistical features. Experiments across summarization, question answering, and instruction-following tasks achieve high teacher identification accuracy. Our results demonstrate that PoS templates constitute a generalizable, low-overhead, and highly informative paradigm for teaching provenance, with significant implications for LLM copyright protection and regulatory compliance auditing.

Detect footprints of teacher LLMs in distillationIdentify teacher models from student outputsInfer teachers violating proprietary model terms

This study investigates the effectiveness boundaries of on-policy distillation in training reasoning models, with a focus on critical factors such as teacher selection, self-distillation context, and token-wise optimal policies. To this end, the authors propose a training-free diagnostic framework that, for the first time, quantifies at the token level the alignment between distillation signals and ideal gradient directions. This is achieved through a combination of ideal node gradient derivation and a scalable directional rollout algorithm to efficiently compute gradient alignment scores (measured via cosine similarity). The analysis reveals that distillation signals are more informative when the student errs but may introduce noise along correct reasoning paths. Furthermore, the optimal distillation configuration is highly dependent on both student capability and task characteristics, with no universally optimal setting.

on-policy distillationper-token supervisionreasoning models

Standard online policy distillation (OPD) often suffers from training instability due to high noise in teacher trajectories and large variance in supervision signals. To address this, this work proposes the BRTS framework, which introduces a novel multi-trajectory sampling and prioritization mechanism that selects high-quality teacher trajectories primarily based on their correctness and secondarily on their alignment with the student’s behavior. Additionally, BRTS incorporates a ground-truth conditioned recovery strategy to handle challenging samples and integrates an auxiliary supervision loss to enhance training stability. Evaluated on demanding mathematical reasoning benchmarks—including AIME 2024/2025 and AMC 2023—BRTS significantly outperforms standard OPD, with the most pronounced gains observed on the hardest problem subsets.

High-VarianceOn-Policy DistillationReasoning

This work addresses the limitations of existing post-training methods, which lack fine-grained reasoning guidance under sparse verifier rewards, and conventional online policy distillation approaches that overlook the interdependencies among multiple rollouts from the same prompt. The authors propose a multi-turn online policy distillation framework that, for the first time, jointly leverages both successful and failed trajectories generated by the student model under identical prompts to construct contrastive teacher signals. This enables conditional, instance-adaptive dense supervision grounded in peer experience. By integrating positive peer imitation with a success-failure contrastive mechanism and modeling multi-trajectory context, the method significantly outperforms standard distillation baselines across programming, mathematical reasoning, scientific question answering, and tool-use tasks, while achieving higher alignment between teacher signals and verifier rewards, thereby validating its efficacy.

multi-rollout learningon-policy distillationpeer conditioning

In strong-to-weak policy distillation, full-trajectory supervised training often suffers from inefficiency due to “local teachability collapse” in later segments. This work formally defines this phenomenon and introduces a trajectory-adaptive supervision release mechanism that dynamically truncates uninformative supervision regions while preserving only the most discriminative portions of teacher feedback for training. The method leverages NLTK sentence segmentation, Top-K candidate margin analysis, and Bayesian Information Criterion (BIC)-based change-point detection to identify effective supervision boundaries. Evaluated across multiple student models in the Qwen3 series, the approach consistently outperforms full-trajectory distillation on five in-domain benchmarks and demonstrates superior generalization on out-of-domain tasks.

dense feedbacklocal teachability collapseon-policy distillation

Latest Papers

What's happening recently
View more

This work addresses the unreliability of teacher supervision signals in on-policy distillation by introducing, for the first time, a prompt-level teacher consistency reliability metric \( R \), and empirically validates its positive correlation with distillation performance. To efficiently estimate \( R \) without extensive teacher inference, the authors propose ROUGE-5 F1 as a proxy metric, enabling prompts to be ranked in descending order of reliability and integrated into a reliability-aware prompt scheduling mechanism. The approach combines independently sampled student trajectories with teacher trajectories filtered by a verifier, achieving consistent and significant improvements over existing baselines across mathematical and code generation tasks on Qwen3 and Gemma4 models, and demonstrating robust gains in all six configurations of FiRe-OPD and ExOPD.

On-Policy DistillationPrompt OrderingPrompt Reliability

Hot Scholars

MS

Minjoon Seo

Config Intelligence; KAIST
Artificial IntelligenceLanguage Modeling
SW

Shilei Wen

bytedance.com
computer visionmachine learning
SP

Shrideep Pallickara

Professor of Computer Science, Colorado State University
Distributed SystemsCyberinfrastructureSpatial Data ScienceClouds
SK

Sungdong Kim

KAIST
Machine LearningNatural Language Processing