Inner Momentum for Differentially Private Muon

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that gradient clipping in differentially private training disrupts the singular vector geometry of the Muon optimizer. To mitigate this, we propose an inner momentum mechanism based on historical gradient averaging, which decouples public scaling from covariance residuals to constrain distortion. Theoretically, we prove that finite Newton-Schulz iterations stably preserve the corrected polar factor and establish a rigorous upper bound on clipping-induced distortion. Empirically, our approach comprehensively outperforms the DP-Muon baseline on GPT-2 fine-tuning tasks, yielding significant improvements in BLEU and ROUGE-L scores. Furthermore, non-private diagnostics demonstrate a 2–4% reduction in pre-noise polar error, validating the efficacy of the proposed geometric correction strategy for privacy-preserving optimization.
📝 Abstract
Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
Problem

Research questions and friction points this paper is trying to address.

Differential Privacy
Muon Optimizer
Gradient Clipping
Singular Vector Distortion
Polar Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Muon Optimizer
Inner Momentum
Gradient Clipping
Newton-Schulz Iteration
🔎 Similar Papers