An extrapolated and provably convergent algorithm for nonlinear matrix decomposition with the ReLU function

๐Ÿ“… 2025-03-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This paper studies nonlinear matrix decomposition with ReLU activation (ReLU-NMD): given a sparse nonnegative matrix (X), find a low-rank matrix (Theta) such that (X approx max(0, Theta)). First, we reveal the ill-posedness of the latent ReLU-NMD model. Second, we propose 3B-ReLU-NMDโ€”a novel reformulation enabling decoupled rank constraints. Third, we design eBCD-NMD, the first extrapolated block coordinate descent algorithm for ReLU-NMD endowed with rigorous convergence guarantees, integrating BCD, extrapolation, and nonconvex optimization analysis. We establish global convergence of eBCD-NMD and prove its accelerated iteration complexity. Empirical evaluations on synthetic and real-world datasets demonstrate that eBCD-NMD consistently outperforms state-of-the-art methods in both reconstruction accuracy and computational efficiency. The method is broadly applicable to tasks including data compression, imputation of non-randomly missing entries, and manifold learning.

Technology Category

Machine Learning: Matrix & Tensor MethodsSearch and Optimization: Non-convex OptimizationData Mining & Knowledge Management: Data Compression

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
๐Ÿ“ Abstract
Nonlinear matrix decomposition (NMD) with the ReLU function, denoted ReLU-NMD, is the following problem: given a sparse, nonnegative matrix $X$ and a factorization rank $r$, identify a rank-$r$ matrix $Theta$ such that $Xapprox max(0,Theta)$. This decomposition finds application in data compression, matrix completion with entries missing not at random, and manifold learning. The standard ReLU-NMD model minimizes the least squares error, that is, $|X - max(0,Theta)|_F^2$. The corresponding optimization problem is nondifferentiable and highly nonconvex. This motivated Saul to propose an alternative model, Latent-ReLU-NMD, where a latent variable $Z$ is introduced and satisfies $max(0,Z)=X$ while minimizing $|Z - Theta|_F^2$ (``A nonlinear matrix decomposition for mining the zeros of sparse data'', SIAM J. Math. Data Sci., 2022). Our first contribution is to show that the two formulations may yield different low-rank solutions $Theta$; in particular, we show that Latent-ReLU-NMD can be ill-posed when ReLU-NMD is not, meaning that there are instances in which the infimum of Latent-ReLU-NMD is not attained while that of ReLU-NMD is. We also consider another alternative model, called 3B-ReLU-NMD, which parameterizes $Theta=WH$, where $W$ has $r$ columns and $H$ has $r$ rows, allowing one to get rid of the rank constraint in Latent-ReLU-NMD. Our second contribution is to prove the convergence of a block coordinate descent (BCD) applied to 3B-ReLU-NMD and referred to as BCD-NMD. Our third contribution is a novel extrapolated variant of BCD-NMD, dubbed eBCD-NMD, which we prove is also convergent under mild assumptions. We illustrate the significant acceleration effect of eBCD-NMD compared to BCD-NMD, and also show that eBCD-NMD performs well against the state of the art on synthetic and real-world data sets.
Problem

Research questions and friction points this paper is trying to address.

Develops convergent algorithm for nonlinear ReLU matrix decomposition
Compares ReLU-NMD and Latent-ReLU-NMD solution differences
Proposes extrapolated BCD method accelerating decomposition convergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extrapolated BCD algorithm for ReLU-NMD
Convergent block coordinate descent method
Latent variable model for matrix decomposition
๐Ÿ’ผ Related Jobs
No related jobs found.
Nicolas Gillis
Nicolas Gillis
University of Mons
optimizationdata sciencenumerical linear algebrasignal processing
Margherita Porcelli
Margherita Porcelli
Associate Professor, University of Florence
MathematicsNumerical AnalysisNonlinear Optimization
G
Giovanni Seraghiti
University of Mons, Rue de Houdain 9, 7000 Mons, Belgium, and Dipartimento di Ingegneria Industriale, Universitร  degli Studi di Firenze, Viale Morgagni 40/44, 50134 Firenze, Italia. Member of the INdAM Research Group GNCS.