What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between data supervision and computational and storage bottlenecks in language model training by proposing a differentiated supervision strategy grounded in a "diagnosis-to-data" principle. Methodologically, it introduces the concept of a "moving bottleneck" to construct a causality-preserving dynamic data scheduling mechanism. Furthermore, it integrates counterfactual availability analysis, paired supervision, and contextual opportunity ranking to optimize the training trajectory. Experiments on 350M-parameter models demonstrate that continuous intervention strategies significantly outperform staged replacement approaches, effectively improving accuracy in long-context question answering. These findings establish a novel paradigm for mitigating resource bottlenecks during model training.
📝 Abstract
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.
Problem

Research questions and friction points this paper is trying to address.

language model training
circuit bottlenecks
data curriculum
supervision strategy
Innovation

Methods, ideas, or system contributions that make the work stand out.

circuit bottleneck
diagnosis-to-data principle
availability counterfactuals
prerequisite ordering
conditional arbitration
🔎 Similar Papers
No similar papers found.