Score
Designs, builds, or analyzes methods that decompose a nonnegative data matrix into nonnegative factor matrices (e.g., basis and activation matrices) to produce parts-based, low-rank representations for dimensionality reduction, pattern discovery, and source separation. This includes choosing or estimating the number of components, imposing constraints or regularization on factors, and using the factors to isolate recurring patterns or separate background from signal.
This paper addresses the identifiability problem of five nonnegative matrix factorization (NMF) models—LBA, LCA, EMA, PLSA, and standard NMF—under a unified probabilistic modeling, convex geometric analysis, and constrained optimization framework. We rigorously prove, for the first time, that the identifiability of solutions to all five models is *exactly equivalent* to that of standard NMF. This equivalence bridges long-standing theoretical fragmentation across cognitive modeling, machine learning, and statistical inference. Our framework provides the first general, cross-model identifiability criterion. Experiments on real-world time-budget data empirically validate the theoretical predictions and clarify boundary conditions relative to prototype analysis and related methods. The core contribution lies in unifying disparate identifiability theories—previously scattered across disciplines—under the foundational NMF paradigm, thereby establishing a rigorous, model-agnostic basis for principled model selection and interpretability assessment.
This work addresses the multi-level low-rank (MLR) matrix approximation problem under the Frobenius norm, tackling three core challenges: hierarchical structural partitioning (row/column stratification), rank allocation (optimizing individual block ranks under a total storage budget), and joint factor fitting. We propose the first end-to-end joint optimization framework for MLR matrices, unifying structural design, rank assignment, and factor learning within a single model. Our approach employs hierarchical block-diagonal parameterization, alternating optimization, and a constrained rank allocation algorithm to achieve coordinated optimization. The resulting approximation preserves matrix-vector multiplication complexity at O(n). Empirical evaluation on multiple benchmark datasets shows that our method reduces approximation error by 35% on average compared to single-level low-rank baselines, significantly improving both accuracy and storage efficiency. The implementation is publicly available.
To address the poor interpretability of Non-negative Matrix Factorization (NMF) in biomedical applications and the limited low-dimensional representation capability of conventional rule-based methods, this paper proposes RuleNMF—a novel non-negative data representation framework that embeds symbolic rules into the NMF architecture. Its core innovation is the first realization of rule-regularized latent factor mapping: numerical latent variables are explicitly encoded as high-coverage, semantically transparent interval- or category-based rule subsets, preserving part-based structure while enabling precise semantic interpretation. By integrating rule-guided constrained optimization and rule–feature alignment modeling, RuleNMF significantly enhances interpretability and downstream performance in multi-label supervised NMF and focused embedding tasks. Moreover, it supports quantitative attribution analysis—e.g., attribute importance scoring—and cross-factor relational reasoning.
This paper addresses the weak theoretical foundations of matrix decomposition in machine learning by systematically constructing a self-consistent, comprehensive, and modern-application-oriented pedagogical framework. Methodologically, it grounds the exposition in numerical linear algebra and matrix analysis, unifying classical decompositions—including LU, QR, SVD, and block triangular factorizations—while integrating numerical stability analysis and Hermitian/Hilbert space theory. Crucially, it bridges traditional numerical analysis with deep learning’s backpropagation setting, emphasizing differentiability and computational robustness of decompositions in algorithm design and model optimization. The primary contribution is a compact, dual-purpose (teaching and research) knowledge system that fills critical gaps in both theoretical coherence and machine-learning relevance present in existing literature, thereby providing rigorous mathematical foundations for high-dimensional data modeling and efficient training.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
This work addresses the limitation of traditional tensor decomposition methods in incorporating prior knowledge from computational models when analyzing high-dimensional multi-way data, such as metabolomics datasets, which hinders the discovery of interpretable patterns. The authors propose a knowledge-guided coupled tensor decomposition framework that, for the first time, jointly analyzes real observational data and simulated data generated by computational models under linear coupling constraints. This approach enhances both robustness and interpretability of extracted patterns in noisy settings and successfully identifies latent inconsistencies between model predictions and empirical observations in real metabolomics data. The results demonstrate the method’s effectiveness and novelty in seamlessly integrating domain-specific prior knowledge with data-driven analysis.
This work addresses the lack of rigorous theoretical guarantees for sparsity-induced identifiability in general real-valued three-factor matrix decompositions. The authors propose a novel decomposition strategy that reformulates the original problem into two coupled auxiliary factorizations. By integrating spectral approximation error analysis, high-probability bound derivations, and structural consistency theory, they establish—for the first time—a rigorous theoretical framework for sparsity-induced identifiability under this setting. Their results elucidate the critical role of sparse coefficients in determining recovery conditions, convergence behavior, and structural preservation. Monte Carlo experiments confirm that sparsity substantially enhances both the recoverability and structural fidelity of the factor matrices, with empirical findings closely aligning with theoretical predictions.
This work investigates the design of pooling-free scattering networks employing fixed monomial nonlinearities to maximize separability for data with low intrinsic dimensionality. By integrating frame theory, geometric measure theory, and moment analysis, the study provides the first geometric characterization of a scattering network’s separation capacity, establishing theoretical bounds for feature extractors operating on low-dimensional rectifiable data. The core contribution consists of two practical design principles: the network’s filters must span a sufficiently broad frequency range, and the frame formed by these filters—when coupled with the data’s geometric structure through a coupling matrix—must exhibit a well-conditioned condition number. These criteria jointly ensure significantly enhanced separation performance, offering concrete guidance for the construction of effective scattering architectures tailored to geometrically structured low-dimensional data.