Predicting Block-Coordinate Performance via Cross-Curvature

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of selecting between synchronous and sequential block update strategies in machine learning optimization, where their relative efficacy depends on objective geometry and iteration count. To resolve this, we propose a unified analytical framework grounded in cross-block curvature that expresses single-step loss differences as curvature functions and derives K-step iterative loss comparisons with explicit error bounds. The predictive accuracy is further validated through higher-order Taylor expansions combined with recursive parameter measurements. Applied to neural networks, federated learning, and LoRA scenarios, the proposed theory correctly identifies the lower-loss strategy in 83.4%–98.0% of cases, improving to 93.2%–100% when integrated with recursive measurement. This work thereby enables rigorous quantitative evaluation of block update strategies.
📝 Abstract
Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an $O(\eta^3)$ remainder, where $\eta$ is the learning rate. We derive a signed loss comparison after $K$ iterations with $O(K\eta^3)$ error under regularity conditions and $\eta K\le T$ for fixed $T$, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0\% of iterations for the neural network, 83.4\% for federated learning, and 97.6\% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0\%, 93.2\%, and 99.6\%, respectively.
Problem

Research questions and friction points this paper is trying to address.

block-coordinate optimization
Jacobi update
Gauss-Seidel update
cross-curvature
performance prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Block-Coordinate Optimization
Cross-Curvature
Jacobi vs Gauss-Seidel
Loss Prediction
Unified Theory
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.