🤖 AI Summary
This study addresses the unclear sample complexity and step-size costs required to translate theoretical directions into effective finite updates in natural gradient optimization. By leveraging exponential family models and Fisher information matrix analysis, the work quantifies worst-case sample complexity under a KL-divergence budget constraint and derives theoretical lower bounds for step-size selection using computational complexity theory. The authors prove that determining an effective step size remains NP-hard even when the exact natural gradient is known, revealing that step-size selection, rather than direction estimation, constitutes the primary computational bottleneck. Furthermore, matching sample complexity lower bounds are established. Experimental results demonstrate that introducing a 10% KL margin in frozen-feature classifier heads yields a joint success rate exceeding 93%.
📝 Abstract
What information is needed to turn a natural-gradient direction into a useful finite update? Under a population Kullback-Leibler (KL) budget, we call a step useful if it is feasible and loses at most a fraction $\varepsilon$ of the best feasible gain along the direction. We construct a four-state exponential family whose laws share their initial gradient, scalar Fisher information and natural gradient, yet two laws have disjoint useful-step sets. With these quantities supplied exactly and the law otherwise known only through draws, the family's worst-case sample complexity is $\Theta(\log(1/\delta)/(p\varepsilon^2))$ for small $\varepsilon$, where $p$ scales rare-state probabilities and $\delta$ is the failure probability. The budget is fixed and the optimal gain stays bounded away from zero, so the step length, not the direction, carries this cost. For succinctly described event-tilt models, returning a useful step is NP-hard even with the exact natural gradient and efficient exact sampling. Recovering the unit natural gradient to constant error is also NP-hard even in a two-parameter logistic family with Fisher condition number at most 3. We also give matching sample bounds for event tilts, sample bounds for damped Fisher solves and a population-KL certificate for affine classifiers. In frozen-feature classifier heads, stopping at a sampled KL boundary succeeds in about half of the trials, and a 10% KL margin raises joint success above 93% at a KL budget of 0.01. Thus, knowing where to move is not enough: how far to move can carry an update's entire cost.