Sequential Functional Structured Tucker Compression for Large Language Model Attentions

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing attention compression methods for large language models, which typically overlook shared structural information and are prone to representation drift. To overcome these issues, this work proposes FTC, a framework that integrates sequential adaptation with functional structured Tucker decomposition to jointly model the query, key, and value head structures. By independently handling output projections to mitigate representation shifts, FTC achieves training-free, gradient-free, and model-state-aware compression. This approach transcends conventional independent matrix approximation paradigms. Extensive experiments on models ranging from 6B to 32B parameters demonstrate that FTC attains the lowest perplexity among competing methods. Notably, under aggressive compression ratios, it significantly outperforms baseline approaches while effectively preserving performance across downstream tasks.
📝 Abstract
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Attention Compression
Post-training Compression
Representation Shift
Structured Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Compression
Structured Tucker Decomposition
Post-training Compression
Large Language Models
Representation Shift
🔎 Similar Papers
No similar papers found.
J
Jiangfeng Chen
University of Manitoba
Xinyu Wang
Xinyu Wang
PhD student, McGill University
Large Language ModelRetrieval Augmented GenerationQuantization
T
Tianshuo Yan
The University of Hong Kong
H
Hanwei Wu
Simpleway
X
Xiao-Wen Chang
McGill University
Y
Yang Zhang
University of Manitoba
L
Lei Ding
University of Manitoba