Score
Designs and implements optimization algorithms and system-level techniques for estimating low-rank matrices or tensor factors and for solving penalized MAP (maximum a posteriori) or regularized objective functions. Builds memory- and computation-efficient update rules, factorized representations, streaming/blocked algorithms, and parallelization strategies so penalized fitting and inference scale to problems with millions of parameters under tight memory and time constraints.
This work addresses efficiency bottlenecks in solving large-scale linear systems and approximating matrix norms. We propose a multilevel randomized sketching preconditioned iterative method, integrating Nyström low-rank approximation, sparse random sketching, and multilevel preconditioning. It establishes the first multilevel sketched preconditioning framework grounded in the natural average condition number. Theoretical contributions include: (1) optimal complexity $ ilde{O}(n^2 + d_lambda^omega)$ for solving regularized linear systems; (2) accelerated complexity $ ilde{O}(n^{2.065} + k^omega)$ for systems with $k$ outlying singular values; and (3) Schatten-$p$ norm approximation—particularly the nuclear norm—at $ ilde{O}(n^{2.11})$, improving upon the prior best $ ilde{O}(n^{2.18})$. These advances significantly enhance computational efficiency for key subproblems in applications such as Gaussian process regression.
In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.
L₀-regularized optimization remains computationally challenging for efficient solving and practical deployment in machine learning, statistical modeling, and signal processing. To address this, we propose the first customizable L₀ modeling framework, integrating high-performance exact solvers based on branch-and-bound, dynamic programming, and sparse optimization. Implemented as a modular Python toolbox, it enables flexible user specification of objective functions and constraints, and provides out-of-the-box machine learning pipelines. Empirical evaluation demonstrates substantial improvements in feature selection accuracy and model interpretability across diverse sparse modeling tasks, achieving state-of-the-art solution quality and runtime efficiency. The framework is open-source, designed for extensibility, and supports seamless integration into industrial-scale machine learning systems.
This work addresses the multi-level low-rank (MLR) matrix approximation problem under the Frobenius norm, tackling three core challenges: hierarchical structural partitioning (row/column stratification), rank allocation (optimizing individual block ranks under a total storage budget), and joint factor fitting. We propose the first end-to-end joint optimization framework for MLR matrices, unifying structural design, rank assignment, and factor learning within a single model. Our approach employs hierarchical block-diagonal parameterization, alternating optimization, and a constrained rank allocation algorithm to achieve coordinated optimization. The resulting approximation preserves matrix-vector multiplication complexity at O(n). Empirical evaluation on multiple benchmark datasets shows that our method reduces approximation error by 35% on average compared to single-level low-rank baselines, significantly improving both accuracy and storage efficiency. The implementation is publicly available.
This work addresses the computational inefficiency of matrix functions—such as square roots, inverse roots, and orthogonalization—in neural network training, which stems from traditional iterative methods’ reliance on prior spectral information and their inability to adapt to dynamically changing matrix spectra. The authors propose PRISM, a novel framework that enables adaptive computation of matrix functions without requiring any prior knowledge of the spectrum. PRISM constructs, at each iteration, a polynomial surrogate of the current spectrum using random sketching and relies predominantly on GPU-friendly matrix multiplications. This approach automatically adapts to spectral shifts during training, substantially reducing computational overhead. When integrated into Shampoo and Muon optimizers, PRISM maintains optimization accuracy while significantly decreasing both iteration counts and wall-clock runtime.
This study addresses the computational challenges posed by high-dimensional regularized estimating equations, which often exhibit non-gradient structures, asymmetric Jacobians, over-identification, non-smoothness, non-convexity, or nested optimization, rendering standard penalized methods inefficient. To tackle this, the paper proposes a unified formulation of such problems as fixed-point equations and systematically develops four computational paradigms—minimization-based, Dantzig-type, regularization-based, and fixed-point-based—integrating strategies from penalized optimization, constrained linear programming, iterative root-finding, and proximal fixed-point iterations. This cohesive framework substantially enhances both solvability and algorithmic stability for high-dimensional regularized estimating equations, demonstrating broad applicability to complex settings such as longitudinal data analysis and survival modeling.
Traditional least squares (LS) is commonly treated as a static fitting tool, making it incompatible with end-to-end differentiable learning frameworks. This work reformulates LS as a differentiable operator: analytical gradients are derived via the implicit function theorem, and matrix-free iterative solvers eliminate explicit storage and inversion of large matrices; a customized backward pass enables implicit incorporation of structural constraints—such as sparsity and conservatism—without modifying the forward computation. The resulting operator integrates seamlessly into automatic differentiation systems and supports joint optimization with neural networks. Experiments demonstrate computational efficiency at the 50-million-parameter scale and successful application to physics-constrained generative modeling and performance-driven hyperparameter optimization in Gaussian processes. By unifying classical LS estimation with modern differentiable programming, our approach significantly enhances the expressivity and practical utility of least squares within contemporary machine learning pipelines.
This project addresses efficient randomized algorithms for three fundamental matrix computations under resource constraints: (1) low-rank positive semidefinite (PSD) matrix approximation, aiming for exact low-rank reconstruction with minimal element accesses; (2) estimation of implicit matrix properties—trace, diagonal entries, and row norms—using only matrix-vector products; and (3) numerically robust floating-point solutions to overdetermined linear least-squares problems. Methodologically, we propose a randomized pivoted Cholesky decomposition to drastically reduce element accesses in PSD approximation; introduce a novel leave-one-out framework for property estimation, achieving optimal sample complexity; and design the first randomized least-squares solver that simultaneously attains asymptotic acceleration—O(nd log d) time for n × d matrices—and numerical backward stability, matching the accuracy of direct methods. These contributions advance both theoretical guarantees and practical efficiency in large-scale, memory- or query-constrained numerical linear algebra.
For composite optimization problems where ill-conditioning or over-parameterization of the smooth mapping impedes subgradient methods to sublinear convergence, this paper proposes a preconditioned Levenberg–Marquardt-type subgradient method. Unlike conventional approaches, it does not require the smooth mapping to be well-conditioned; instead, it only assumes standard regularity conditions on the convex component, thereby achieving global linear convergence—even in over-parameterized regimes where traditional methods suffer from convergence degradation. The algorithm integrates preconditioning with adaptive regularization, fully exploiting the composite structure of the objective. Theoretical analysis establishes its applicability to canonical nonconvex problems, including phase retrieval, matrix sensing, and tensor decomposition. Numerical experiments demonstrate significantly accelerated convergence and enhanced robustness.
This work addresses the critical challenge that modern GPU-accelerated linear programming solvers—such as cuPDLP, which is based on the primal-dual hybrid gradient (PDHG) algorithm—exhibit performance highly sensitive to hyperparameters, yet lack tuning methods with provable generalization guarantees. For the first time, this study establishes structural relationships between hyperparameters and solution trajectories for multiple adaptive techniques in complex first-order LP solvers, including preconditioning, restart strategies, and smoothed weight updates. By integrating convergence analysis of PDHG with a model of structural sensitivity, the authors propose a data-driven hyperparameter learning framework that offers theoretical generalization guarantees under polynomial sample complexity. Experimental results demonstrate that the framework significantly enhances solver efficiency across diverse problem instances.