SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scaling local learning to billion-parameter pre-training and the memory bottlenecks inherent in backpropagation. To this end, we propose Shared Output Local Learning (SOLL), a method that introduces read-only shared readout heads to propagate deep-layer information, thereby eliminating the update locking problem. By integrating pipeline parallelism, SOLL overcomes the limitations of conventional private readouts, achieving efficient local learning-based pre-training for billion-scale Transformer models for the first time. Experimental results demonstrate that SOLL attains an accuracy comparable to standard backpropagation while improving throughput by 1.44× and significantly reducing memory consumption. These findings validate both the feasibility and the efficiency advantages of SOLL for large-scale pre-training.
📝 Abstract
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.
Problem

Research questions and friction points this paper is trying to address.

Local Learning
Update Locking
Large Language Models
Pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local Learning
Shared-Output
Billion-Parameter Pretraining
Update Locking
Throughput Optimization
🔎 Similar Papers
No similar papers found.