The cost of useful natural gradient updates

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear sample complexity and step-size costs required to translate theoretical directions into effective finite updates in natural gradient optimization. By leveraging exponential family models and Fisher information matrix analysis, the work quantifies worst-case sample complexity under a KL-divergence budget constraint and derives theoretical lower bounds for step-size selection using computational complexity theory. The authors prove that determining an effective step size remains NP-hard even when the exact natural gradient is known, revealing that step-size selection, rather than direction estimation, constitutes the primary computational bottleneck. Furthermore, matching sample complexity lower bounds are established. Experimental results demonstrate that introducing a 10% KL margin in frozen-feature classifier heads yields a joint success rate exceeding 93%.
📝 Abstract
What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family whose laws share their initial gradient, scalar Fisher information and natural gradient, yet two laws have disjoint useful-step sets. With these quantities supplied exactly and the law otherwise known only through draws, the family's worst-case sample complexity is $\Theta(\log(1/\delta)/(p\varepsilon^2))$ for small $\varepsilon$, where $p$ scales rare-state probabilities and $\delta$ is the failure probability. The budget is fixed and the optimal gain stays bounded away from zero, so the step length, not the direction, carries this cost. For succinctly described event-tilt models, returning a useful step is NP-hard even with the exact natural gradient and efficient exact sampling. Recovering the unit natural gradient to constant error is also NP-hard even in a two-parameter logistic family with Fisher condition number at most 3. We also give matching sample bounds for event tilts, sample bounds for damped Fisher solves and a population-KL certificate for affine classifiers. In frozen-feature classifier heads, stopping at a sampled KL boundary succeeds in about half of the trials, and a 10% KL margin raises joint success above 93% at a KL budget of 0.01. Thus, knowing where to move is not enough: how far to move can carry an update's entire cost.
Problem

Research questions and friction points this paper is trying to address.

natural gradient
step size
sample complexity
KL divergence budget
computational complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

natural gradient
sample complexity
KL budget
NP-hard
Fisher information
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Subhransu S. Bhattacharjee
School of Computing, The Australian National University
Dylan Campbell
Dylan Campbell
Lecturer, Australian National University
RegistrationGlobal optimization3D Reconstruction3D/Stereo Scene Analysis
R
Rahul Shome
School of Computing, The Australian National University