Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the functional degradation in existing tensor network compression methods, which minimize weight approximation error while neglecting activation distributions. We propose AWT, a training-free, activation-aware diagonal preconditioning calibration method that optimizes Tensor-Train (TT) and Tree Tensor Network (TTN) decompositions without modifying the underlying solvers. Functioning as a modular plugin compatible with post-training quantization techniques, AWT substantially enhances the functional fidelity of compressed large language models. Extensive experiments demonstrate that AWT significantly reduces perplexity and improves downstream task performance across diverse LLMs, effectively narrowing the performance gap between compressed models and their dense baselines.
πŸ“ Abstract
Post-training tensor-network compression replaces Transformer linear layers with Tensor Train (TT) or Tree Tensor Network (TTN) operators, but standard decompositions minimize weight-space Frobenius error rather than functional error under the layer's activation distribution. We propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling. Across Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B, AWT consistently improves vanilla TT/TTN tensorization at 2-6 times compression: under single-operator replacement, AWT closes 12-35% of the WikiText perplexity gap to the dense baseline across the three model families and 2-6 times compression settings; while under multi-operator Llama suffix replacement it closes 27-60% across attention-group and all-seven-matrix settings. The gains also transfer to downstream HellaSwag and ARC-Challenge evaluations. We further show that diagonal preconditioning is a robustness-modularity tradeoff rather than a diagonal-covariance assumption: a dense full-covariance oracle wins its own weighted objective in 80/81 cases, yet diagonal AWT gives better held-out functional fidelity in 53/81 cases. Together, these results position AWT as a principled, modular preconditioner for improving functional fidelity in fixed TT/TTN compression pipelines without modifying the decomposition solver.
Problem

Research questions and friction points this paper is trying to address.

tensor-network compression
large language models
functional error
weight tensorization
post-training compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation-Aware Weight Tensorization
Tensor-Network Compression
Training-Free Calibration
Diagonal Preconditioning
Functional Fidelity