QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the redundant KV loading and computational overhead in sparse attention caused by independent query processing. We propose an optimization framework that jointly processes adjacent queries to reuse shared KV entries. The core innovation, Shift-Compare Set Decomposition (SCSD), transforms irregular set operations into data-parallel primitives, enabling cascaded sharing and tile-aware execution. This is further integrated with a workload-aware sparse attention mechanism and pipeline optimization strategies. Experimental results demonstrate that our approach reduces kernel latency by 55.1% and KV data throughput by 55.9%, while decreasing time-to-first-token by 36.8%, all with negligible degradation in model accuracy.
📝 Abstract
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead. Cascaded sharing captures reuse hierarchically at multiple granularities. A tile-aware execution strategy balances sharing granularity with hardware tile utilization and selectively removes low-importance query-specific tails to eliminate underutilized tiles. We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism. Compared with the state-of-the-art sparse-attention kernel, QUILT reduces average kernel latency by up to 55.1% and processed KV data by up to 55.9%, while reducing time-to-first-token (TTFT) latency by up to 36.8% with negligible accuracy degradation.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Long-context Prefill
KV Cache Reuse
Memory Traffic
Redundant Computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Shared Query Execution
Set Decomposition
KV Reuse
Prefill Optimization
🔎 Similar Papers
No similar papers found.
Z
Zhenduo Zhao
Huawei Technologies Co., Ltd.
Q
Qihui Zhou
Huawei Technologies Co., Ltd.
Mingcong Song
Mingcong Song
University of Florida
computer architecture/systemmachine learning
Z
Zhiyi Chen
Huawei Technologies Co., Ltd.
C
Chuangguan Ye
Huawei Technologies Co., Ltd.
F
Fengfan Hou
Huawei Technologies Co., Ltd.
Z
Zequn Gong
Huawei Technologies Co., Ltd.
J
Jing Li
Huawei Technologies Co., Ltd.
H
Hongjie Si
Huawei Technologies Co., Ltd.
G
Guoping Long
Huawei Technologies Co., Ltd.