🤖 AI Summary
This study elucidates the mechanisms and boundaries of the weak-to-strong generalization phenomenon in random feature networks. Leveraging random matrix theory, gradient flow analysis, and the Gaussian universality hypothesis, it derives deterministic equivalents for teacher-student model errors while characterizing asymptotic behaviors under ReLU activation and the effects of varying stopping times. The core contributions are threefold: first, proving that quadratic error improvement attains the theoretical lower bound; second, precisely delineating the transition regime for non-quadratic improvement; and third, establishing an exact scaling law whereby student error scales as the square of teacher error. Collectively, this work systematically completes the theoretical framework for weak-to-strong generalization.
📝 Abstract
Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.