🤖 AI Summary
This study addresses the dual bottlenecks of end-to-end backpropagation: prohibitive memory overhead and incompatibility with biologically plausible local learning mechanisms. We propose an efficient training framework that integrates I-JEPA pretraining, neurally localized weight updates, and pruning-based signal transmission, enabling flexible trade-offs between accuracy and memory consumption by modulating gradient propagation spans. By replacing global error backpropagation with local learning rules, the proposed approach substantially reduces GPU memory footprint. Evaluated on benchmarks including ImageNet, our method achieves competitive classification performance with minimal peak GPU memory usage. This work establishes a novel paradigm for low-resource training of large-scale models that is both biologically grounded and practically viable for real-world engineering deployment.
📝 Abstract
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly localized. Therefore, we propose LocalProp, a training procedure that locally updates the weights of a model. Our neuro-localized weight updates follow the"pre-training then fine-tuning"paradigm, where the pre-training is based on I-JEPA. After locally updating the weights, a pruning operation is performed, followed by a short final fine-tuning phase. Pruning helps by sending the learning signal from higher blocks to lower blocks. We perform experiments on several datasets, including large-scale benchmarks such as ImageNet, and empirically show that LocalProp reaches good performance at a fraction of GPU peak memory. By varying the number of jointly optimized blocks, we identify gradient-propagation span as a practical control over the accuracy-memory trade-off.