Score
Designs and applies quantitative analyses and estimators for the effective rank of matrices and the spectrum of Hessians, builds methods to estimate matrix rank and curvature spectra from data, and analyzes how stochastic noise scales with curvature; uses these tools to relate curvature (Hessian) eigenvalue structure to optimization behaviour and to predict scaling phenomena such as slowdown that are independent of raw parameter count.
Accurate estimation of the Hessian spectrum for large-scale foundation models has long been hindered by prohibitive computational costs, with existing approaches relying on small models or strong structural approximations. This work proposes a sharded local finite difference method compatible with Fully Sharded Data Parallel (FSDP), combined with stochastic Lanczos quadrature, enabling the first efficient and accurate computation of the true Hessian spectrum on hundred-billion-parameter language models. The method incurs only constant-level training overhead and reveals that widely used block-diagonal approximations severely break down even in moderately sized LLMs. Furthermore, it establishes a scalable paradigm for Hessian spectral analysis, supporting numerical stability studies under fp32/bf16 precision and end-to-end scaling law modeling.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
This work investigates the intrinsic relationship between SGD training dynamics and the spectral structure of empirical Hessian and gradient matrices in high-dimensional multiclass classification. Methodologically, it employs theoretical modeling and spectral analysis to characterize the evolution of these matrices throughout training. The key contribution is the first rigorous proof that, in deep multilayer networks, SGD trajectories align layerwise with the anomalous eigensubspaces—i.e., those spanned by top eigenvalues—of both the layer-wise Hessian and gradient matrices. Crucially, the rank of the final-layer anomalous subspace monotonically degenerates during optimization and becomes severely deficient upon convergence to a suboptimal classifier, serving as a spectral diagnostic for suboptimal convergence in overparameterized networks. This result is formally established for high-dimensional mixture models and single-/two-layer neural networks. Moreover, the work quantifies the direct link between the evolution of the final-layer anomalous subspace and classification accuracy, offering a novel spectral-geometric perspective on deep learning optimization dynamics.
This work addresses stochastic optimization on the generalized Stiefel manifold—defined as the set of matrices $X$ satisfying $X^ op B X = I_p$—which arises in canonical correlation analysis (CCA), independent component analysis (ICA), and generalized eigenvalue problems (GEVP). Conventional Riemannian approaches require exact knowledge of $B$ and enforce the constraint via costly retraction or projection at each iteration. We propose the first projection-free stochastic algorithm: it operates using only unbiased stochastic estimates of $B$, employs a constrained-stable stochastic gradient update dominated by matrix multiplications, and converges almost surely to constrained critical points. Theoretically, its convergence rate matches that of full-information Riemannian methods under standard assumptions. Empirically, each iteration incurs significantly lower computational cost while achieving comparable accuracy.
This work addresses the efficient and robust approximation of the Hessian (or its inverse) in stochastic optimization, unifying the formulation across Euclidean space, the symmetric positive-definite (SPD) manifold, and general Lie groups. We establish, for the first time, that under mild conditions the Hessian-fitting objective on a Lie group is strongly convex—transforming the original ill-conditioned problem into a well-conditioned Lie-group optimization task. Leveraging this property, we propose an efficient preconditioner construction method based on Hessian-vector products and sparse structural priors. We further conduct a systematic geometric analysis comparing the adaptability and convergence behavior of second-order and adaptive methods—including BFGS, natural gradient, and PSGD—within this unified framework. Theoretical analysis and empirical evaluation demonstrate that our approach significantly improves the efficiency, numerical stability, and scalability of second-order information utilization in large-scale stochastic optimization.
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.
This work addresses the suboptimality of standard stochastic gradient descent (SGD) convergence analyses, which neglect the geometric heterogeneity of gradient noise in parameter space. To remedy this, the authors propose Curvature-Weighted Gradient Diversity (CWGD), a novel noise metric that weights sample-wise gradient diversity by the inverse square root of the Hessian, thereby aligning the noise characterization with the underlying optimization geometry. Building on this metric, they design the CWGD-Cosine learning rate schedule. Theoretical analysis demonstrates that this approach reduces the asymptotic optimization error to half that of standard cosine annealing. Empirical validation across varying condition numbers, batch sizes, and noise structures consistently shows approximately 20% lower final error on average, with negligible computational overhead.
This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.
This work addresses the gap between abstract Riemannian geometry and practical algorithmic implementation by systematically developing a computationally tractable geometric framework for Riemannian optimization. Focusing on canonical matrix manifolds—Stiefel, Grassmann, and symmetric positive-definite (SPD) manifolds—it explicitly derives core geometric structures, including tangent spaces, metric tensors, Levi-Civita connections, curvature operators, and geodesics, all expressed in coordinate- and matrix-based forms amenable to numerical computation. Furthermore, it provides closed-form expressions for the Riemannian gradient, Hessian, exponential map, and retraction operators. To the best of our knowledge, this is the first unified formulation that translates classical differential-geometric constructions into a consistent, implementation-ready framework, thereby bridging theory and practice and offering a rigorous foundation for efficient and accurate algorithm design in Riemannian optimization and geometric machine learning.
This work investigates how the sharpness of neural network loss landscapes in classification tasks is influenced by the distribution of data across classes. By analytically deriving the Hessian spectrum of linear networks under mean squared error loss for arbitrary width, depth, and dataset size, and corroborating the findings with large-scale numerical experiments, the study establishes—for the first time under minimal simplifying assumptions—an exact analytical relationship between the largest eigenvalue of the Hessian and the proportion of samples in the dominant class. The theory demonstrates that solution sharpness is governed primarily by the prevalence of the most frequent class and exhibits remarkable robustness across diverse nonlinear architectures and realistic settings, thereby systematically elucidating the critical role of data structure in shaping optimization landscapes.