QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottleneck caused by the linear growth of Multi-head Latent Attention (MLA) cache with context length, as well as the error amplification inherent in dual-path quantization. We propose a function-aligned dual-path quantization framework that achieves the first joint INT4 quantization of content and RoPE caches. By establishing a dual-path error model and designing tailored objective functions, combined with an offline-fusable transformation space, attention output reconstruction, and native low-bit compute kernels, our method enables efficient KV cache compression. Experimental results demonstrate that under a 128K context length, the approach attains a 3.59× physical compression ratio and improves serving throughput by 5.168× over BF16 baselines with negligible accuracy degradation.
📝 Abstract
Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

Multi-Head Latent Attention
KV cache quantization
low-bit caching
RoPE path error
memory efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Head Latent Attention
Low-Bit Quantization
KV Cache Compression
Function-Aligned Dual-Path
Attention Kernel
🔎 Similar Papers
Z
Zunhai Su
The University of Hong Kong
Y
Yuxuan Sun
Meituan LongCat Team
Jianchao Tan
Jianchao Tan
Meituan
LLMAutomated Machine LearningComputer GraphicsComputer Vision
T
Tao Zhang
South China University of Technology
R
Ruihan Hu
Harbin Institute of Technology
Y
Yuchen Xie
Meituan LongCat Team
X
Xunliang Cai
Meituan LongCat Team
N
Ngai Wong
The University of Hong Kong