🤖 AI Summary
This study addresses the scalability limitations of subspace clustering imposed by dense self-expression matrices and full-affinity spectral clustering. We propose LoomSC, a framework that jointly learns latent features and self-representations through projector factorization and exact spectral reduction, thereby circumventing the construction of complete affinity matrices. By introducing alternating Procrustes updates to preserve orthogonality and combining implicit coefficient matrices with non-negative quadratic affinities, LoomSC achieves linear time and space complexity. Experimental results demonstrate that LoomSC surpasses the strongest baseline by an average accuracy margin of 6.66% across five image benchmarks. Notably, it maintains over 99.8% accuracy even when scaling to 500,000 samples, effectively balancing computational efficiency with superior clustering performance.
📝 Abstract
Dense self-expression matrices and full-affinity spectral clustering limit the scalability of subspace clustering. We introduce the Latent Orthogonal Optimization Model for Subspace Clustering (LoomSC), a framework that addresses both bottlenecks through projector factorization and exact spectral reduction. Motivated by the spectral structure of least-squares regression, LoomSC jointly learns latent features and a projector self-representation through two thin factors. Alternating Procrustes and least-squares updates preserve the sample factor's orthogonality while keeping the coefficient matrix implicit. We construct a nonnegative quadratic affinity that preserves the projector's support. An exact feature map then reduces its normalized spectral problem to an eigenproblem whose dimension depends only on the factor width. Neither the full affinity nor the sample Laplacian needs to be formed. Our analysis quantifies the projector approximation and identifies conditions for subspace preservation and within-subspace connectivity. For fixed dimensions and iteration budgets, the complete pipeline has linear time and memory complexity in the number of samples. Across five image-clustering benchmarks, LoomSC ranks first or second in all 15 dataset-metric comparisons against 9 state-of-the-art baselines. Its mean accuracy exceeds the highest baseline mean by 6.66 percentage points. Synthetic experiments scale to 500,000 samples while maintaining at least 99.8% accuracy.