student-teacher training

Design and implement training systems in which a teacher model produces targets or constraints (e.g., soft labels, pseudo-labels, consistency targets) that are used to supervise and shape a student model's learning. This includes choosing teacher update rules (mean/EMA teachers), formulating teacher–student loss functions and auxiliary regularizers, propagating pseudo-labels, allocating per-example supervision budgets, and optimizing student parameters under teacher guidance.

student-teachertraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Guiding Through Complexity: What Makes Good Supervision for Hard Reasoning Tasks?

Oct 27, 2024
XH
Xuan He
🏛️ Tsinghua University | University of California, Los Angeles

This work investigates how to effectively enhance large language models’ performance on high-difficulty reasoning tasks—such as mathematical proof generation and complex logical deduction—using weak supervision sources (e.g., non-expert annotators or off-the-shelf AI systems). Addressing the trade-off between supervision quality and task difficulty, we establish a key empirical finding: supervision for hard tasks with high step-level error rates yields superior model performance compared to error-free supervision on easy tasks; critically, step-level error rate proves a more informative training signal than final-answer accuracy. Building on this insight, we propose a collaborative supervision paradigm that jointly leverages subtasks and hard tasks, integrated with supervised data mixing, dynamic reweighting, and empirically grounded evaluation. Our approach achieves up to 30% absolute accuracy gains on challenging benchmarks such as MATH. All code and datasets are publicly released.

Effective supervision by weak teacher modelsImpact of step-wise error ratesImproving LLMs on hard reasoning tasks

Modelling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting

Oct 21, 2024
RT
Robin Th'eriault
🏛️ Scuola Normale Superiore di Pisa | The University of British Columbia | University of Bologna

This work investigates the learning mechanism of restricted Boltzmann machines (RBMs) for structured data within a teacher–student framework, specifically addressing whether a student RBM can recover the true latent representation from data generated by a teacher RBM exhibiting correlations among hidden variables. Using statistical-physics-based mean-field analysis and temperature-regularized inference, we systematically characterize how structural strength—quantified by the number of latent patterns and inter-pattern correlation in weight rows—affects the critical sample size required for successful learning. We find that enhanced structure drastically reduces the necessary sample complexity; in the absence of correlations, performance is independent of pattern count; and excessively low inference temperatures suppress pattern acquisition, leading to learning failure. Crucially, we establish for the first time that the student can achieve exact one-to-one or one-to-many pattern matching—surpassing the conventional two-hidden-unit limitation. These results provide the first analytically tractable generative-model foundation for the “lottery ticket hypothesis.”

Analyzing impact of pattern correlations on critical data requirementsExploring temperature effects on teacher pattern learnabilityStudying RBM learning of structured data in teacher-student setting

Reinforcement Teaching

Apr 25, 2022
AL
Alex Lewandowski
🏛️ University of Alberta | Huawei Technologies Canada Co., Ltd. | Google Brain

Existing meta-learning methods suffer from limited generalizability, often being confined to specific algorithms or requiring differentiability assumptions. This paper proposes a general reinforcement learning–driven meta-learning framework that trains a teacher policy to dynamically guide arbitrary student algorithms—without imposing structural or differentiability constraints on the student. Key contributions include: (i) the first unified pedagogical paradigm for meta-learning; (ii) a parameter-behavior encoder that implicitly infers the student’s internal parameter state from its input-output behavior; and (iii) a reward function grounded in learning progress. Experiments across supervised and reinforcement learning tasks demonstrate that our framework significantly outperforms baselines relying on heuristic rewards and handcrafted state representations, validating its broad generalizability and empirical effectiveness.

AdaptabilityMachine Learning EfficiencyMeta-Learning

A Pattern Language for Machine Learning Tasks

Jul 02, 2024
BR
Benjamin Rodatz
🏛️ Compositional Intelligence | Quantinuum | University of Oxford

Existing machine learning frameworks suffer from insufficient formalization of objective functions and lack a unified, cross-domain behavioral design paradigm. Method: We propose an equation-constrained compositional function modeling approach for learners, constructing task graphs and compositional semantic graphs to enable model-agnostic behavioral specification and optimization. We introduce a novel task-oriented pattern language framework and the “manipulator” task paradigm, supporting end-to-end, architecture-agnostic, and adversarial-training-free minimal editing of data attributes. Contribution/Results: Theoretically, our work integrates formal methods and theoretical computer science principles. Empirically, we demonstrate precise, controllable, and interpretable behavioral editing on small-scale models under stable training—without stochastic sampling or data intervention—yielding significant improvements in deployment efficiency and formal verifiability.

Creating model-agnostic tasks for stable small-scale ML modelsDeveloping a graphical mathematics for unified ML task designFormalizing objective functions as equality constraints on learners

Representational Alignment Supports Effective Machine Teaching

Jun 06, 2024
IS
Ilia Sucholutsky
🏛️ Princeton University | University of Cambridge | Stevens Institute of Technology | MPI | Anthropic | NYU | MIT | University College London | UC Berkeley | The Alan Turing Institute

Existing machine learning teaching frameworks overlook representational alignment between teachers and students, prioritizing only model accuracy. Method: We propose GRADE, a representation-alignment-driven pedagogical optimization framework. Leveraging controlled machine–machine and machine–human teaching experiments, we formally define and quantify the relationship between representational alignment and teaching utility, introducing the alignment-driven teaching utility curve. We further design GRADE-Match, a cross-modal teacher–student matching algorithm that optimizes representational adaptation. Contribution/Results: Experiments demonstrate that improved representational alignment significantly enhances student task accuracy—moderated by class size and representation diversity. In simulated teaching settings, GRADE-Match achieves an average 12.3% improvement in learning outcomes. GRADE establishes a novel paradigm for interpretable, optimization-aware intelligent teaching systems grounded in representational alignment principles.

Characterize teacher expertise and student learningOptimize student-teacher matching with GRADEStudy pedagogy and representational alignment

Latest Papers

What's happening recently
View more

This work addresses the distribution mismatch commonly faced by large language model agents in supervised fine-tuning, where training relies on complete teacher demonstrations while testing depends on student-generated contexts. The authors formulate online policy data construction as a budget allocation problem and propose replacing lengthy or costly filtered teacher trajectories with a small number of unfiltered, short-step teacher continuations, strategically injected into student-induced critical contexts. By systematically exploring the design space of rollout policies, switching time distributions, continuation lengths, and filtering rules—and incorporating a dual-cost model accounting for both teacher inference and supervision signal retention—the method demonstrates strong empirical performance on HotpotQA, ALFWorld, and Terminal-Bench-Dev. Notably, it matches or exceeds existing critical-context filtering baselines on the first two benchmarks at lower computational cost, indicating that a few well-placed teacher steps can substantially enhance training efficiency.

cost-efficient supervisiondistribution mismatchon-policy data augmentation

This study addresses the challenge in knowledge distillation where varying data budgets shift the optimal teacher capacity, complicating effective sample selection. By analyzing relational ranking and score geometry, this work reveals for the first time why smaller teachers excel under low-data regimes and proposes the DVA framework. Employing a small teacher as a proxy, DVA models relational diversity through difficulty filtering and class-conditional volume maximization, establishing a training-free dynamic data selection mechanism that jointly optimizes difficulty matching and signal diversity. Without requiring training-dependent dynamic statistics, the proposed approach achieves performance comparable to dynamic methods while consistently outperforming static baselines, offering a novel paradigm for efficient knowledge distillation.

Data PruningData SelectionKnowledge Distillation

This study addresses a critical gap in machine learning education: the overreliance on pre-labeled datasets, which often obscures the subjectivity and ambiguity inherent in data annotation, leading students to place undue trust in model outputs. To counter this, the authors introduce an innovative pedagogical intervention that transforms manual annotation into an active learning tool. Students annotated hair coverage in skin lesion images using a three-point scale, followed by structured reflections via questionnaires. A cross-institutional experiment involving 43 participants from Fontys University of Applied Sciences (Netherlands) and the IT University of Copenhagen (Denmark) demonstrated that this approach significantly enhanced learners’ awareness of annotation ambiguity, dataset biases, and model limitations. Most participants acknowledged the influence of personal interpretation on labeling decisions and reported higher engagement compared to traditional instruction. This work provides the first empirical evidence supporting subjective annotation as an effective strategy for cultivating critical thinking about AI systems.

biasdata annotationinterpretive diversity

This study addresses the counterproductive effects of fine-tuning large language model (LLM) tutoring systems on pedagogical adaptability metrics, which inadvertently leads to behavioral rigidity and degraded instructional quality. We propose the theoretical hypothesis that independent decision-making averaged over metrics tends to converge toward repetitive optimal solutions, and establish design principles for AI tutor benchmarks accordingly. An empirical investigation is conducted integrating automated scoring, open-source model fine-tuning, blinded expert evaluation, and weight analysis. Our findings reveal a critical over-optimization effect wherein improvements in metric scores are accompanied by significant declines in expert assessments, demonstrating that measurement validity does not entail optimization validity. This work provides essential caveats and actionable design guidelines for deploying LLMs in educational applications, cautioning against uncritical reliance on proxy metrics during alignment.

LLM tutoringmeasurement validitymetric optimization

Hot Scholars

XB

Xiang Bai

Huazhong University of Science and Technology (HUST)
Computer VisionOCR
UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing
TO

Takahiro Ogawa

Hokkaido University
Multimedia ProcessingAIIoTBig Data Analysis
GL

Guang Li

Assistant Professor, Hokkaido University
Dataset DistillationSelf-Supervised LearningData-Centric AIMedical Image Analysis
YQ

Yanmin Qian

Professor, Shanghai Jiao Tong University
Speech and Language ProcessingSignal ProcessingMachine Learning