Clean: Second-order LLM Training at Linear Memory Cost via Nystr\"om Sketching

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing memory-efficient optimizers discard curvature information, while full-curvature methods incur prohibitive memory overhead. To bridge this gap, this work proposes Clean, an optimizer that leverages randomized Nyström approximation, SOAP preconditioning, and low-rank reconstruction to reduce the memory complexity of second-order optimization from quadratic to linear while fully preserving curvature information. Additionally, a quantized variant, Q-Clean, is introduced to balance efficiency and performance. Experimental results demonstrate that Clean accelerates convergence by 26% compared to AdamW and reduces memory consumption by over 50% relative to Muon. Notably, it enables the pretraining of 13B-parameter models on a single GPU for the first time.
📝 Abstract
Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right preconditioners in SOAP, and to reduce the optimizer's memory complexity from quadratic to linear in terms of model dimensions. We subsequently reintegrate the off-subspace components to capture curvature information beyond the low-rank approximation, preserving rich curvature at minimal memory cost. We further propose Q-Clean, a low-precision variant that aggressively compresses optimizer states. Q-Clean reduces optimizer memory consumption by \textbf{over 50\%} compared to Muon when pre-training a LLaMA-1.3B architecture, all while maintaining strong and competitive predictive performance. Notably, Clean operates with a smaller optimizer-state footprint than standard AdamW while reaching AdamW's final performance \textbf{26\% faster} in wall-clock time. Furthermore, our methods uniquely enable the pre-training of a 13B-parameter model on a single 80GB GPU, providing a scalable, efficient, and accessible approach to large-scale model optimization.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Second-order Optimization
Memory Efficiency
Curvature Information
Optimizer States
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nyström Sketching
Second-order Optimization
Memory-efficient Optimizer
Low-precision Training
Large Language Models
🔎 Similar Papers
No similar papers found.