An Active-Bottleneck Mechanism for Weak-to-Strong Generalization

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the theoretical conditions under which a student model surpasses its teacher in weak-to-strong generalization without explicit regularization. Based on two-stage linear regression and random feature models, the analysis employs ridge regression, spectral techniques, and power-law covariance modeling to propose an "active bottleneck" principle. This principle reveals that a finite number of pseudo-labels and limited model width serve as implicit regularization mechanisms that effectively filter out teacher noise. The work derives closed-form threshold intervals under which the student outperforms the teacher, precisely characterizing both unified and split improvement regimes. These findings provide a rigorous theoretical foundation for understanding enhanced generalization under weak supervision.
📝 Abstract
Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from $n$ labeled examples and a student is trained solely on the teacher's predictions on $m$ fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when $m$ lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of $m$. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for $m$ in a bounded interval, and one where it occurs only once the student width $N_S$ exceeds an explicit threshold. Both regimes are governed by a single"active-bottleneck"principle: whichever of $m$ or $N_S$ is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
Problem

Research questions and friction points this paper is trying to address.

weak-to-strong generalization
two-stage learning
pseudo-labels
random-feature model
active-bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weak-to-Strong Generalization
Active-Bottleneck Mechanism
Two-Stage Learning
Implicit Regularization
Random Feature Model
💼 Related Jobs
No related jobs found.
M
Mohammad Zeinalpour
Department of Computer Engineering, Sharif University of Technology
Amir Najafi
Amir Najafi
imec, Belgium
SOC designUltra-low-power on-chip communicationEnergy-efficient architectures