Score
Designs and analyzes algorithms and procedures that decompose matrices into factor matrices under explicit sparsity constraints, producing low‑rank sparse components or basis vectors (including for embedding matrices) that can be used to isolate concept directions, enable precise projection‑based editing, and support targeted erasure or modification modules.
This work addresses the challenge of unifying diverse structured sparsity patterns for efficient model compression and acceleration. The authors propose S³, an algebraic framework that formally integrates three core components—View (tensor reshaping), Block (atomic pruning units), and Scope (sparsity decision range)—to express a wide spectrum of sparsity patterns, ranging from fine-grained N:M sparsity to coarse-grained channel pruning, within a single formalism. Notably, S³ enables cross-tensor collaborative sparsification. Building upon this framework, the authors incorporate Optimal Brain Damage and Surgeon algorithms to develop structured variants of OBS/OBD. These methods significantly outperform current state-of-the-art second-order heuristic approaches in terms of output reconstruction accuracy.
Structured random matrices lack a unified theoretical analysis and a general design framework in randomized linear algebra. Method: This paper introduces the “Oblivious Subspace Injection” (OSI) property, establishing the first decoupled abstract analytical framework that separates correctness proofs of algorithms from instantiation-specific verification. Contribution/Results: We prove that sparse random matrices, random triangular transforms, and tensor-product-structured matrices all satisfy OSI, thereby unifying their dimensionality-reduction fidelity guarantees for tasks such as low-rank approximation and least-squares regression. Leveraging this framework, we design accelerated algorithms with near-optimal time complexity. Empirical evaluation on synthetic datasets and scientific computing benchmarks confirms both efficiency and practical utility.
This work addresses the multi-level low-rank (MLR) matrix approximation problem under the Frobenius norm, tackling three core challenges: hierarchical structural partitioning (row/column stratification), rank allocation (optimizing individual block ranks under a total storage budget), and joint factor fitting. We propose the first end-to-end joint optimization framework for MLR matrices, unifying structural design, rank assignment, and factor learning within a single model. Our approach employs hierarchical block-diagonal parameterization, alternating optimization, and a constrained rank allocation algorithm to achieve coordinated optimization. The resulting approximation preserves matrix-vector multiplication complexity at O(n). Empirical evaluation on multiple benchmark datasets shows that our method reduces approximation error by 35% on average compared to single-level low-rank baselines, significantly improving both accuracy and storage efficiency. The implementation is publicly available.
Computing exact leverage scores for Kronecker-product structured matrices in large-scale least-squares problems is computationally prohibitive, while existing approximation methods incur statistical bias and high overhead. Method: We propose the first efficient exact leverage score algorithm tailored to Kronecker-structured matrices. Leveraging the inherent tensor structure, our method designs a near-linear-time framework for exact leverage score computation and sampling—bypassing costly full-matrix SVD or biased sketching approximations. Contribution/Results: Theoretically and empirically, our algorithm achieves significantly lower sampling error than state-of-the-art approximate methods (e.g., FJLT- or CountSketch-accelerated approaches), while maintaining substantially lower time complexity than full SVD. This work establishes the first scalable, exact, and efficient leverage score sampling scheme for Kronecker-structured matrices, enabling improved structured random projections and large-scale regression.
Existing sparse matrix optimization methods often coarsely classify structured sparsity (e.g., clustered non-zeros) as either fully dense or fully sparse, leading to redundant zero computations in fixed-block formats (e.g., BCSR) or substantial overhead in variable-block approaches due to unknown loop bounds at compile time. This work proposes a region-aware, multi-stage compilation framework that automatically identifies high-benefit variable-size blocks, statically infers dynamic loop bounds, and generates customized vectorized code—balancing efficiency and adaptability. Key techniques include sparse partition analysis, domain-specific code generation, loop vectorization, and compile-time scheduling specialization. Evaluated on the SuiteSparse dataset, our approach achieves 1.07×, 2.73×, and 1.9× higher single-threaded SpMV performance over Intel MKL, CSR5, and Partial Strided Codelets, respectively; parallel scalability further enhances throughput.
This work addresses the lack of rigorous theoretical guarantees for sparsity-induced identifiability in general real-valued three-factor matrix decompositions. The authors propose a novel decomposition strategy that reformulates the original problem into two coupled auxiliary factorizations. By integrating spectral approximation error analysis, high-probability bound derivations, and structural consistency theory, they establish—for the first time—a rigorous theoretical framework for sparsity-induced identifiability under this setting. Their results elucidate the critical role of sparse coefficients in determining recovery conditions, convergence behavior, and structural preservation. Monte Carlo experiments confirm that sparsity substantially enhances both the recoverability and structural fidelity of the factor matrices, with empirical findings closely aligning with theoretical predictions.
This work addresses the problem of efficiently constructing an approximate matrix with a prescribed binary sparsity pattern using only black-box access via matrix-vector product queries. The authors introduce the degeneracy of the sparsity pattern as a unified and tight measure of query complexity, overcoming limitations inherent in traditional graph coloring approaches. Leveraging this notion, they propose an adaptive querying strategy together with a polynomial-time algorithm that achieves a near-optimal approximation using only Õ(degen(S)) queries while avoiding computational bottlenecks. Furthermore, they establish an information-theoretic lower bound of Ω(degen(S)) on the query complexity for any sparsity pattern S, thereby proving the optimality of their approach.
This work addresses the high computational cost of sparse matrix ordering in linear systems arising from triangular meshes. We propose an efficient ordering algorithm that accelerates nested dissection and rapidly constructs elimination trees by moderately relaxing constraints on partition balance and optimality, while integrating local block ordering with a separator-based quotient graph compression strategy. The method innovatively trades a controlled degradation in ordering quality for substantial computational speedup, preserving the fill-reducing structure required for Cholesky factorization while bypassing its most expensive phases. When integrated into commercial CPU/GPU sparse Cholesky solvers, our approach significantly reduces ordering time in graphics applications and achieves up to a 6.27× improvement in overall solver performance.