Output-aware Residual Stream Pruning for Large Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation in existing residual stream pruning methods caused by neglecting downstream task sensitivity. To overcome this limitation, we propose an output-aware residual stream pruning framework that quantifies the impact of dimensional perturbations on model outputs via a second-order approximation of the KL divergence. By introducing a sensitivity-weighted covariance matrix, the method reformulates subspace selection as an eigendecomposition problem, moving beyond the reliance on activation reconstruction error alone. This approach effectively balances accuracy and efficiency while substantially reducing calibration KL divergence and improving both perplexity and downstream task performance. Overall, these results validate the effectiveness of modeling output propagation mechanisms for efficient model compression.
📝 Abstract
Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.
Problem

Research questions and friction points this paper is trying to address.

Residual Stream Pruning
Large Language Models
Model Compression
Sensitivity-aware
Inference Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Stream Pruning
Sensitivity-aware Pruning
Output KL Divergence
Spectral Upper Bound
Large Language Models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chayne Thrash
Department of Computer Science, Vanderbilt University
K
Kevin Chen
Department of Computer Science, Vanderbilt University
Soheil Kolouri
Soheil Kolouri
Computer Science, Vanderbilt University, Nashville, TN
Machine LearningOptimal TransportComputer Vision