Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究针对现有量化方法依赖标量敏感度代理导致的精度损失问题,提出了一种基于激活意识和跨层优化的新量化方法CASA。
📝 Abstract
Mixed-precision weight quantization is commonly formulated as a Multiple-Choice Knapsack Problem (MCKP), yet existing solvers rely on scalar sensitivity proxies that collapse each weight matrix's Hessian into a single number and treat every module independently. We prove that even the optimal scalar proxy incurs multiplicative distortion up to $\sqrt{κ(\mathbf{A})κ(\mathbf{B})}$ relative to the full activation-aware quadratic, where $κ(\mathbf{A})$ and $κ(\mathbf{B})$ denote the condition numbers of the input- and output-side Hessian factors. This bound varies from $10^1$ to $10^{13}$ for typical LLM modules, making inter-module sensitivity ranking unreliable. To address these limitations, we propose Cross-layer Activation-aware Sensitivity Allocation (CASA), a two-phase method. In Stage 1, the scalar proxy is replaced by an activation-aware metric derived from the Kronecker-factored Hessian, reducing the MCKP to a form whose continuous relaxation admits a closed-form solution. In Stage 2, a cross-layer-aware local search evaluates bit-width updates using the end-to-end model loss. Experiments on multiple LLMs across different bit budgets show that CASA achieves lower perplexity than the latest scalar-proxy baselines, especially at ultra-low bit-widths ($<3$ bits per weight). Moreover, the performance gain in zero-shot accuracy tracks the per-model average condition-number over modules, confirming the distortion bound as a practical indicator of scalar-proxy failure.
Problem

Research questions and friction points this paper is trying to address.

Mixed-precision Quantization
Scalar Sensitivity
Hessian Matrix
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-layer Activation-aware Sensitivity Allocation
Mixed-precision Quantization
Kronecker-factored Hessian
Large Language Models
Ultra-low Bit-widths
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.