🤖 AI Summary
This study addresses the mismatch between data supervision and computational and storage bottlenecks in language model training by proposing a differentiated supervision strategy grounded in a "diagnosis-to-data" principle. Methodologically, it introduces the concept of a "moving bottleneck" to construct a causality-preserving dynamic data scheduling mechanism. Furthermore, it integrates counterfactual availability analysis, paired supervision, and contextual opportunity ranking to optimize the training trajectory. Experiments on 350M-parameter models demonstrate that continuous intervention strategies significantly outperform staged replacement approaches, effectively improving accuracy in long-context question answering. These findings establish a novel paradigm for mitigating resource bottlenecks during model training.
📝 Abstract
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.