Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

📅 2026-09-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对低秩LLM压缩问题,提出了一种三级优化方法,包括矩阵级、块级和模型级优化,以减少独立压缩矩阵在非线性前向传递中的误差累积。
📝 Abstract
Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened SVD~(L1), block-level joint optimization~(L2), and end-to-end language-modeling loss refinement~(L3), all from 256 calibration sequences, with no instruction or recovery data. On LLaMA-7B at 60% compression, the chain reduces WikiText-2 perplexity from 42.1 to 19.1 to 11.4. The block-level stage acts as a regularizer: skipping it worsens Penn Treebank (PTB) perplexity by 24 points, a gap that additional end-to-end training did not close in our experiments. Perplexity gains hold across 20-80% compression, five architectures up to 13B parameters, and both in-distribution and out-of-distribution benchmarks, though the cross-architecture rows use architecture-specific configurations and the ratio sweep was not run under one common protocol. With more calibration data, skipping the block-level stage becomes competitive, revealing an offline compute--data trade-off. We therefore claim improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.
Problem

Research questions and friction points this paper is trying to address.

low-rank LLM compression
per-matrix SVD truncation
error compounding
nonlinear forward pass
Innovation

Methods, ideas, or system contributions that make the work stand out.

three-level optimization
whitened SVD
block-level joint optimization
end-to-end language-modeling loss refinement
low-rank LLM compression
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huicheng Zhang
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
X
Xiyao Feng
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
Z
Ze-Tong Li
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
Chengkai Zhu
Chengkai Zhu
Ph.D. Student, The Hong Kong University of Science and Technology (Guangzhou)
Quantum information
X
Xiao Shi
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
X
Xiwei Pan
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
Jinguo Liu
Jinguo Liu
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
G
Ge Bai
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China
X
Xin Wang
Hong Kong University of Science and Technology (Guangzhou), Guangdong 511453, China