Low-Rank Compression of Pretrained Models via Randomized Subspace Iteration

📅 2026-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying large-scale pretrained models under high computational and memory costs. Existing low-rank compression methods, such as randomized SVD, often fail to preserve approximation accuracy when the singular value spectrum of weight matrices decays slowly, leading to degraded predictive performance. By analyzing the perturbation in softmax outputs induced by low-rank compression, this study establishes a theoretical connection between approximation error and classification performance. To mitigate this issue, the authors propose employing randomized subspace iteration with multiple power iterations to enhance spectral gap separation and improve approximation quality. The method consistently outperforms conventional randomized SVD across both convolutional and Transformer architectures, achieving near-optimal low-rank approximations under aggressive compression while maintaining higher prediction accuracy.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Large Vision ModelsData Mining & Knowledge Management: Data Compression

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
The massive scale of pretrained models has made efficient compression essential for practical deployment. Low-rank decomposition based on the singular value decomposition (SVD) provides a principled approach for model reduction, but its exact computation is expensive for large weight matrices. Randomized alternatives such as randomized SVD (RSVD) improve efficiency, yet they can suffer from poor approximation quality when the singular value spectrum decays slowly, a regime commonly observed in modern pretrained models. In this work, we address this limitation from both theoretical and empirical perspectives. First, we establish a connection between low-rank approximation error and predictive performance by analyzing softmax perturbations, showing that deviations in class probabilities are controlled by the spectral error of the compressed weights. Second, we demonstrate that RSVD is inadequate, and we propose randomized subspace iteration (RSI) as a more effective alternative. By incorporating multiple power iterations, RSI improves spectral separation and provides a controllable mechanism for enhancing approximation quality. We evaluate our approach on both convolutional networks and transformer-based architectures. Our results show that RSI achieves near-optimal approximation quality while outperforming RSVD in predictive accuracy under aggressive compression, enabling efficient model compression.
Problem

Research questions and friction points this paper is trying to address.

low-rank compression
pretrained models
randomized SVD
singular value spectrum
approximation quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Randomized Subspace Iteration
Low-Rank Compression
Spectral Approximation
Pretrained Models
Model Compression
💼 Related Jobs
No related jobs found.
F
Farhad Pourkamali-Anaraki
Department of Mathematical and Statistical Sciences, University of Colorado Denver