HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决长上下文和多轮交互中的高效预填充、紧凑KV缓存存储及准确检索问题,提出HySparse2架构,采用两级KV共享方法。
📝 Abstract
Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architecture with two-level KV sharing. At the outer level, KV Bridging adopts a YOCO-style self-decoder and cross-decoder structure, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention (SWA), while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner level, HySparse2 retains HySparse's core KV Reuse design with two refinements. First, it replaces block-level sparsity with token-level sparsity for finer long-context retrieval. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing allows all cross-decoder KV caches to be constructed from self-decoder hidden states. Prefill can therefore exit after the self-decoder, skipping all cross-decoder layers. On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.
Problem

Research questions and friction points this paper is trying to address.

long-horizon
multi-turn agents
efficient prefill
compact KV-cache storage
accurate long-context retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Sparse Attention
Two-Level KV Sharing
Token-Level Sparsity
KV Reuse
Sliding Window Attention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jianyu Wei
Jianyu Wei
USTC & MSRA Joint PhD
LLM InfraInference SystemQuantizationKernelCo-design
Y
Yizhao Gao
LLM-Core, Xiaomi
Q
Qihao Zhang
LLM-Core, Xiaomi
S
Shimao Chen
LLM-Core, Xiaomi
Zhengju Tang
Zhengju Tang
Peking University
Y
Yu Cheng
LLM-Core, Xiaomi
S
Shengjie Zhou
LLM-Core, Xiaomi
Zihan Jiang
Zihan Jiang
Huawei
AI BenchmarkingDistributed Deep LearningWorkload Characterization.
Y
Yifan Song
LLM-Core, Xiaomi
H
Hailin Zhang
LLM-Core, Xiaomi
L
Liang Zhao
LLM-Core, Xiaomi
B
Bo Yang
LLM-Core, Xiaomi
G
Gang Wang
LLM-Core, Xiaomi
Shijie Cao
Shijie Cao
Microsoft Research Asia
Efficient Deep LearningDeep Learning SystemComputer Architecture
F
Fuli Luo
LLM-Core, Xiaomi