How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ( \sqrt{ \log d } ) Divergence

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper quantifies the efficiency loss incurred when natural gradient descent (NGD) deviates from the Fisher–Rao shortest geodesic. By defining an “inefficiency ratio” and establishing tensor-based criteria, we combine differential geometry with information theory to analyze the geometric overhead of NGD across various parameter families and identify its optimality conditions. Theoretically, we prove that in scale-product families the inefficiency ratio grows as Θ(√log d) with dimensionality, revealing a high-dimensional degradation mechanism. Empirically, we validate three scenarios: zero overhead under quadratic potentials, finite upper bounds for bounded skewness, and pronounced degradation in high-dimensional scale families. This work provides a rigorous geometric framework for understanding the suboptimality and computational costs of NGD.
📝 Abstract
Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \(R \ge 1\), the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension \(d\). A tensor criterion identifies the regime (I) families, with \(R=1\) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher--Rao diameter yields a dimension-independent bound; and (III) for products of scale families---including Gaussian covariances and Gamma rates---\(R\) grows as \(Θ(\sqrt{\log d})\), unbounded in \(d\). Under a per-step Fisher-chord budget, \(R\) translates to a practical computational cost: NGD requires asymptotically at least \(R\) times as many steps as an optimizer following the Fisher--Rao geodesic. Experiments confirm all three regimes: \(R=1\) to machine precision for quadratic-potential families (I), the categorical bound \(π/(2\sqrt{2})\) is approached but not attained (II), and sampled scale-product \(R\) grows with \(d\), reaching \(R \approx 1.5\) for long, high-dimensional moves (III).
Problem

Research questions and friction points this paper is trying to address.

Natural Gradient Descent
Inefficiency Ratio
Fisher-Rao Distance
Mixture Geodesic
Computational Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Natural Gradient Descent
Inefficiency Ratio
Fisher-Rao Geodesic
Information Geometry
Computational Complexity