Making Analog Training Scale: Co-Designing Mapping, Optimizer, and Converters

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scalability limitations of analog in-memory computing for training modern deep models caused by hardware non-idealities. To overcome this challenge, we propose a mixed-precision training paradigm based on gradient accumulation sensitivity. Through system-algorithm co-design, our approach jointly optimizes weight mapping strategies, preconditioned optimizers, and converter dynamic ranges, while introducing a threshold-triggered pulsing mechanism to enhance training stability under electrochemical RAM simulations. This method successfully scales Transformer training to 123 million parameters, achieving validation loss scaling behavior comparable to digital training. Ultimately, this work provides a scalable solution for large-scale analog training, demonstrating that carefully designed algorithmic-hardware co-optimization can effectively mitigate inherent device imperfections and bridge the performance gap between analog and conventional digital paradigms.
📝 Abstract
Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains an open challenge due to severe hardware non-idealities, including physical weights with finite dynamic range and write granularity, analog-digital converters with finite resolution, and noisy and asymmetric updates. Guided by the insight that gradient accumulation is sensitive to precision and rounding errors, we adopt a mixed-precision training paradigm: executing forward and backward matrix multiplications in the analog domain while computing weight gradients in the digital domain. To enable scalable training, we present a holistic system-algorithm co-design that co-optimizes weight mapping to ensure well-conditioned physical and logical weight profiles, couples a preconditioned optimizer with threshold-triggered open-loop pulsing to stabilize training trajectories, and aligns converter dynamic ranges to suppress quantization errors. Evaluated via hardware-calibrated architectural simulations calibrated with electrochemical RAM measurements, our framework scales Transformer training up to $123\text{M}$ parameters with validation loss scaling as $L\propto N^{-0.231}$, where $N$ is the parameter count, comparable to $L\propto N^{-0.238}$ for digital training.
Problem

Research questions and friction points this paper is trying to address.

Analog in-memory computing
hardware non-idealities
scalable training
deep learning
mixed-precision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analog In-Memory Computing
Mixed-Precision Training
System-Algorithm Co-Design
Preconditioned Optimizer
Transformer Scaling
🔎 Similar Papers
No similar papers found.