scalable low-rank optimization

Designs and implements optimization algorithms and system-level techniques for estimating low-rank matrices or tensor factors and for solving penalized MAP (maximum a posteriori) or regularized objective functions. Builds memory- and computation-efficient update rules, factorized representations, streaming/blocked algorithms, and parallelization strategies so penalized fitting and inference scale to problems with millions of parameters under tight memory and time constraints.

scalablelow-rankoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$230K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Faster Linear Systems and Matrix Norm Approximation via Multi-level Sketched Preconditioning

May 09, 2024
MD
Michal Derezi'nski
🏛️ University of Michigan | New York University

This work addresses efficiency bottlenecks in solving large-scale linear systems and approximating matrix norms. We propose a multilevel randomized sketching preconditioned iterative method, integrating Nyström low-rank approximation, sparse random sketching, and multilevel preconditioning. It establishes the first multilevel sketched preconditioning framework grounded in the natural average condition number. Theoretical contributions include: (1) optimal complexity $ ilde{O}(n^2 + d_lambda^omega)$ for solving regularized linear systems; (2) accelerated complexity $ ilde{O}(n^{2.065} + k^omega)$ for systems with $k$ outlying singular values; and (3) Schatten-$p$ norm approximation—particularly the nuclear norm—at $ ilde{O}(n^{2.11})$, improving upon the prior best $ ilde{O}(n^{2.18})$. These advances significantly enhance computational efficiency for key subproblems in applications such as Gaussian process regression.

Improving algorithms for matrix norm approximationSolving linear systems with outlying singular values efficientlySpeeding up regularized linear systems for semidefinite matrices

Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes

May 28, 2024
JA
Jihao Andreas Lin
🏛️ University of Cambridge | MPI for Intelligent Systems

In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.

Gaussian ProcessesHyperparameter OptimizationLarge-scale Datasets

El0ps: An Exact L0-regularized Problems Solver

Jun 04, 2025
TG
Théo Guyard
🏛️ Polytechnique Montréal | Ensai | Université de Rennes | CentraleSupélec

L₀-regularized optimization remains computationally challenging for efficient solving and practical deployment in machine learning, statistical modeling, and signal processing. To address this, we propose the first customizable L₀ modeling framework, integrating high-performance exact solvers based on branch-and-bound, dynamic programming, and sparse optimization. Implemented as a modular Python toolbox, it enables flexible user specification of objective functions and constraints, and provides out-of-the-box machine learning pipelines. Empirical evaluation demonstrates substantial improvements in feature selection accuracy and model interpretability across diverse sparse modeling tasks, achieving state-of-the-art solution quality and runtime efficiency. The framework is open-source, designed for extensibility, and supports seamless integration into industrial-scale machine learning systems.

Offers state-of-the-art solver for practical applicationsProvides a flexible framework for custom problem instancesSolves L0-regularized problems in machine learning

This work addresses the multi-level low-rank (MLR) matrix approximation problem under the Frobenius norm, tackling three core challenges: hierarchical structural partitioning (row/column stratification), rank allocation (optimizing individual block ranks under a total storage budget), and joint factor fitting. We propose the first end-to-end joint optimization framework for MLR matrices, unifying structural design, rank assignment, and factor learning within a single model. Our approach employs hierarchical block-diagonal parameterization, alternating optimization, and a constrained rank allocation algorithm to achieve coordinated optimization. The resulting approximation preserves matrix-vector multiplication complexity at O(n). Empirical evaluation on multiple benchmark datasets shows that our method reduces approximation error by 35% on average compared to single-level low-rank baselines, significantly improving both accuracy and storage efficiency. The implementation is publicly available.

Allocating block ranks under total storage constraintsOptimizing factor adjustments in multilevel low rank matricesSelecting hierarchical partitions with corresponding ranks and factors

This work addresses the computational inefficiency of matrix functions—such as square roots, inverse roots, and orthogonalization—in neural network training, which stems from traditional iterative methods’ reliance on prior spectral information and their inability to adapt to dynamically changing matrix spectra. The authors propose PRISM, a novel framework that enables adaptive computation of matrix functions without requiring any prior knowledge of the spectrum. PRISM constructs, at each iteration, a polynomial surrogate of the current spectrum using random sketching and relies predominantly on GPU-friendly matrix multiplications. This approach automatically adapts to spectral shifts during training, substantially reducing computational overhead. When integrated into Shampoo and Muon optimizers, PRISM maintains optimization accuracy while significantly decreasing both iteration counts and wall-clock runtime.

adaptive computationdistribution-freematrix functions

Latest Papers

What's happening recently
View more

This study addresses the computational challenges posed by high-dimensional regularized estimating equations, which often exhibit non-gradient structures, asymmetric Jacobians, over-identification, non-smoothness, non-convexity, or nested optimization, rendering standard penalized methods inefficient. To tackle this, the paper proposes a unified formulation of such problems as fixed-point equations and systematically develops four computational paradigms—minimization-based, Dantzig-type, regularization-based, and fixed-point-based—integrating strategies from penalized optimization, constrained linear programming, iterative root-finding, and proximal fixed-point iterations. This cohesive framework substantially enhances both solvability and algorithmic stability for high-dimensional regularized estimating equations, demonstrating broad applicability to complex settings such as longitudinal data analysis and survival modeling.

computational challengesestimating equationshigh-dimensional statistics

Matrix-Free Least Squares Solvers: Values, Gradients, and What to Do With Them

Oct 22, 2025
HR
Hrittik Roy
🏛️ Technical University of Denmark

Traditional least squares (LS) is commonly treated as a static fitting tool, making it incompatible with end-to-end differentiable learning frameworks. This work reformulates LS as a differentiable operator: analytical gradients are derived via the implicit function theorem, and matrix-free iterative solvers eliminate explicit storage and inversion of large matrices; a customized backward pass enables implicit incorporation of structural constraints—such as sparsity and conservatism—without modifying the forward computation. The resulting operator integrates seamlessly into automatic differentiation systems and supports joint optimization with neural networks. Experiments demonstrate computational efficiency at the 50-million-parameter scale and successful application to physics-constrained generative modeling and performance-driven hyperparameter optimization in Gaussian processes. By unifying classical LS estimation with modern differentiable programming, our approach significantly enhances the expressivity and practical utility of least squares within contemporary machine learning pipelines.

Developing differentiable least squares solvers as neural layersEnabling scalable sparse models and constrained optimization applicationsUnlocking least squares potential in modern machine learning

Make the most of what you have: Resource-efficient randomized algorithms for matrix computations

Dec 17, 2025
EN
Ethan N. Epperly
🏛️ California Institute of Technology

This project addresses efficient randomized algorithms for three fundamental matrix computations under resource constraints: (1) low-rank positive semidefinite (PSD) matrix approximation, aiming for exact low-rank reconstruction with minimal element accesses; (2) estimation of implicit matrix properties—trace, diagonal entries, and row norms—using only matrix-vector products; and (3) numerically robust floating-point solutions to overdetermined linear least-squares problems. Methodologically, we propose a randomized pivoted Cholesky decomposition to drastically reduce element accesses in PSD approximation; introduce a novel leave-one-out framework for property estimation, achieving optimal sample complexity; and design the first randomized least-squares solver that simultaneously attains asymptotic acceleration—O(nd log d) time for n × d matrices—and numerical backward stability, matching the accuracy of direct methods. These contributions advance both theoretical guarantees and practical efficiency in large-scale, memory- or query-constrained numerical linear algebra.

Designing resource-efficient randomized algorithms for matrix computations.Developing accurate low-rank approximations with minimal matrix entry access.Ensuring backward stability in randomized linear least squares algorithms.

Preconditioned subgradient method for composite optimization: overparameterization and fast convergence

Sep 14, 2025
MD
Mateo Díaz
🏛️ Johns Hopkins University | Purdue University

For composite optimization problems where ill-conditioning or over-parameterization of the smooth mapping impedes subgradient methods to sublinear convergence, this paper proposes a preconditioned Levenberg–Marquardt-type subgradient method. Unlike conventional approaches, it does not require the smooth mapping to be well-conditioned; instead, it only assumes standard regularity conditions on the convex component, thereby achieving global linear convergence—even in over-parameterized regimes where traditional methods suffer from convergence degradation. The algorithm integrates preconditioning with adaptive regularization, fully exploiting the composite structure of the objective. Theoretical analysis establishes its applicability to canonical nonconvex problems, including phase retrieval, matrix sensing, and tensor decomposition. Numerical experiments demonstrate significantly accelerated convergence and enhanced robustness.

Addressing slow sublinear convergence in overparameterized problemsImproving convergence rates for data science applicationsOptimizing composite functions with ill-conditioned smooth maps

This work addresses the critical challenge that modern GPU-accelerated linear programming solvers—such as cuPDLP, which is based on the primal-dual hybrid gradient (PDHG) algorithm—exhibit performance highly sensitive to hyperparameters, yet lack tuning methods with provable generalization guarantees. For the first time, this study establishes structural relationships between hyperparameters and solution trajectories for multiple adaptive techniques in complex first-order LP solvers, including preconditioning, restart strategies, and smoothed weight updates. By integrating convergence analysis of PDHG with a model of structural sensitivity, the authors propose a data-driven hyperparameter learning framework that offers theoretical generalization guarantees under polynomial sample complexity. Experimental results demonstrate that the framework significantly enhances solver efficiency across diverse problem instances.

first-order methodsgeneralization guaranteesGPU acceleration

Hot Scholars

GX

Gongjun Xu

University of Michigan
StatisticsMachine LearningPsychometrics
YA

Yarden As

ETH Zürich
Artificial IntelligenceReinforcement LearningRobotics
QL

Qihang Lin

The University of Iowa
Continuous optimizationStochastic OptimizationMachine LearningMarkov Decision Process
FD

Fragkiskos D. Malliaros

Professor, CentraleSupélec, Université Paris-Saclay
Graph MiningGraph Machine LearningGraph Neural NetworksNetwork Science
DP

Debdeep Pati

Professor, Department of Statistics, University of Wisconsin - Madison
Bayesian nonparametricshigh-dimensional data analysis