Score
Designs, builds, or analyzes regularizers, optimization procedures, and proximal operators that penalize the sum of singular values (the nuclear/trace norm) of matrices to induce low-rank solutions via singular-value shrinkage and thresholding. This includes constructing objectives and algorithms and proving properties such as effective-rank collapse and scale-invariant subgradient behavior for matrix estimation and recovery.
Low-rank regularization (LRR) optimization is hindered by the non-convexity and discontinuity of the rank function, as well as the non-differentiability and high computational cost of singular value decomposition (SVD). This paper proposes an efficient, fully differentiable, SVD-free generalized low-rank regularization framework that unifies nuclear norm, Schatten-*p* norm, and various non-convex relaxations. Our method leverages matrix power series expansion coupled with random projection to yield a differentiable rank estimator, backed by theoretical convergence guarantees—bias and variance decay rapidly with sample size and iteration count. Implemented via GPU-friendly tensor operations, it seamlessly integrates with arbitrary gradient-based optimizers. Empirical evaluation on matrix completion, multi-task learning, and neural network compression demonstrates 3–8× speedup over SVD-based approaches, while matching or surpassing their accuracy.
This work addresses efficiency bottlenecks in solving large-scale linear systems and approximating matrix norms. We propose a multilevel randomized sketching preconditioned iterative method, integrating Nyström low-rank approximation, sparse random sketching, and multilevel preconditioning. It establishes the first multilevel sketched preconditioning framework grounded in the natural average condition number. Theoretical contributions include: (1) optimal complexity $ ilde{O}(n^2 + d_lambda^omega)$ for solving regularized linear systems; (2) accelerated complexity $ ilde{O}(n^{2.065} + k^omega)$ for systems with $k$ outlying singular values; and (3) Schatten-$p$ norm approximation—particularly the nuclear norm—at $ ilde{O}(n^{2.11})$, improving upon the prior best $ ilde{O}(n^{2.18})$. These advances significantly enhance computational efficiency for key subproblems in applications such as Gaussian process regression.
This work addresses the global optimization challenge of overparameterized nonconvex low-rank matrix recovery under noise, where existing methods often get trapped at non-strict saddle points and lack rigorous theoretical guarantees. We propose a unified analytical framework for escape directions and nonexistence proofs of counterexamples, establishing—for the first time under the Restricted Isometry Property (RIP)—a sharp min-max optimal recovery bound for nearly second-order stationary points in the overparameterized regime. By integrating symmetric and asymmetric parameterizations, balanced regularization design, and higher-order optimization theory, we prove that such stationary points achieve recovery accuracy within a constant factor of convex methods. Moreover, the resulting error bound is sharp with respect to both noise level and solution accuracy.
This work addresses the bottleneck in nonconvex matrix sensing where sample complexity scales quadratically with rank (O(r²)). We propose an efficient factorization-based gradient descent method to recover low-rank positive semidefinite matrices from Gaussian linear measurements. By integrating spectral initialization with a novel probabilistic decoupling analysis, we achieve—for the first time in a nonconvex framework—a sample complexity with only linear rank dependence, O(r), matching the optimal rate of nuclear norm minimization. Theoretically, under Ω(rdκ²) measurements, our algorithm converges linearly to the ground-truth matrix. This result substantially improves the sample efficiency and theoretical guarantees of nonconvex optimization for matrix sensing, breaking the long-standing O(r²) rank-squared barrier inherent in prior nonconvex approaches.
In high-dimensional linear regression with strongly anisotropic covariates and true coefficients aligned with top eigendirections of the covariance matrix, conventional ℓ₂-shrinkage estimators (e.g., ridge regression) can underperform the minimum-ℓ₂-norm interpolator. This work demonstrates that *moderate amplification*—scaling the minimum-norm interpolator by a constant greater than one—substantially reduces generalization error, challenging the long-held belief that shrinkage universally dominates interpolation. We establish the first tight upper and lower bounds on the generalization error under diverging aspect ratios (n/d → 0) for anisotropic covariance structures, proving that amplified interpolation achieves the optimal statistical rate. Our theoretical analysis, combining data splitting and Gaussian random projection techniques, rigorously characterizes the mechanism and limits of amplification. Empirical results confirm that the amplified interpolator consistently outperforms classical regularized estimators on real-world anisotropic data.
To address the suboptimal solutions in low-rank matrix completion (LRMC) caused by nuclear norm’s excessive shrinkage of singular values, this paper proposes a compact nonconvex rank surrogate—the reweighted logarithmic norm (RLogN)—and introduces it to LRMC modeling for the first time. RLogN more accurately approximates matrix rank and mitigates singular value bias. We further develop an alternating direction method of multipliers (ADMM)-based optimization framework with theoretically guaranteed convergence, balancing computational efficiency and numerical stability. On image inpainting tasks, our method significantly outperforms state-of-the-art LRMC algorithms: PSNR improves by 2.1 dB and SSIM by 0.032, yielding more natural visual results and sharper edge recovery. The core contributions are: (1) introducing RLogN as a novel nonconvex rank approximation; (2) designing an efficient, provably convergent solver; and (3) empirically validating its superiority on real-world low-rank prior tasks.
Under scale-invariant normalization mechanisms such as RMSNorm, conventional L2 weight decay struggles to promote generalization in grokking, as it acts only along the radial direction of the weight space. This work proposes Low-Rank Decay (LRD), a spectral regularization method based on the nuclear norm that continuously compresses weight singular values through tangential components, thereby driving structural simplification of representations. For the first time, the geometric dynamics of nuclear norm regularization are introduced into grokking research, revealing that LRD achieves effective rank collapse near low-rank manifolds via a “needle-to-fan” subdifferential expansion. On modular arithmetic tasks, LRD significantly accelerates rank collapse in Query and Key matrices and substantially extends the data regime boundary at which grokking occurs, demonstrating its efficacy in facilitating delayed generalization.
This study investigates whether stochastic rounding (SR) retains its regularizing effect in matrices with constant aspect ratios and examines its impact on the singular value spectrum. By integrating singular value analysis, a stochastic rounding quantization model, and spectral theory, the work demonstrates for the first time that SR not only enhances the smallest singular value but also collectively elevates multiple singular values in the tail of the spectrum. This finding reveals that the regularizing influence of SR extends beyond extreme aspect ratio regimes. The results establish SR as a universal spectral regularization mechanism, thereby broadening its theoretical foundation and application potential in numerical computation and low-precision machine learning.