Probe-Space Preconditioning for Fast and Stable Zero-Order Training

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow convergence of zeroth-order optimization, which limits its competitiveness with backpropagation despite its memory efficiency. We propose 1.5-SPSA, a method that constructs a diagonal preconditioner using only a single forward pass to effectively suppress interference from high-curvature directions, thereby significantly accelerating convergence in the probe space. Furthermore, we achieve efficient implementation by integrating 8-bit packed random generation, Triton fused kernels, and distributed parallelism. Experimental results demonstrate that our approach stably trains the OPT-30B model on commodity GPUs, surpassing baseline accuracy on SST-2 while substantially reducing the required convergence steps. This work establishes a new paradigm for low-resource training of large language models.
📝 Abstract
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes with many perturbations (or probes) but fewer steps, allows 1SPSA (Spall, 1992) to outperform zero order methods like MeZO (Malladi et al., 2023) with less training compute. Next, we introduce 1.5-SPSA, adding a single "clean" forward-pass per step to 1SPSA to calculate a cheap diagonal preconditioner in probe-space, which improves convergence rate and convergence by down-weighting high curvature directions. Benchmarking on 6 post-training datasets on both Qwen3 and OPT model families, we show that 1.5-SPSA achieves State-of-the-Art results over previous ZOO solvers with much less optimization steps. For example, we train OPT-13B (for direct comparison to MeZO) and find 1.5-SPSA achieves +3.1% accuracy on SST-2 over both MeZO and BP in only 70 steps vs. MeZO's 100,000 steps. Finally, we combine an 8-bit-packing random generator, triton fused unpack/apply kernels, and distributed parallelism to achieve fast and stable training of models as large as OPT-30B in-place on commodity GPUs (e.g. A100).
Problem

Research questions and friction points this paper is trying to address.

Zero-Order Optimization
Memory Efficiency
Convergence
Deep Learning
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Order Optimization
Probe-Space Preconditioning
1.5-SPSA
Memory-Efficient Training
Diagonal Preconditioner