LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead of redundant indexer selection and layer-wise key cache redundancy in sparse attention mechanisms. To mitigate these issues, this work proposes an optimization framework that integrates cross-layer shared latent representations with layer-specific scoring. The core innovation lies in replacing discrete token selection with shared continuous representations, incorporating critical decoder absorption into queries, constructing a shared latent cache, and introducing a hierarchical selection algorithm. This design enables training-free calibration to effectively balance generation quality with computational efficiency. Experimental results demonstrate that the proposed method reduces logical index cache storage by 61.1%, improves attention recall by 3.28 percentage points, and accelerates decoding speed by 2.72×.
📝 Abstract
Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing principle of Multi-head Latent Attention across indexer layers. Each layer group constructs a shared latent cache from its first layer's hidden states, while layer-specific scoring enables independent token selection. Absorbing key decoders into queries enables direct scoring of the shared cache without reconstructing historical per-layer keys. We develop training-free calibration and investigate a training-aware instantiation of this principle. To balance quality and computation, a hierarchical selection (HS) variant lets followers independently refine a shared candidate set proposed by the anchor. With four-layer sharing, LatentIndex reduces logical indexer-cache storage by 61.1% on DeepSeek-V3.2. Across DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. HS further achieves 2.30-2.72 times decode indexer speedups over DSA across 8K-128K contexts, retaining most of LatentIndex's recall. LatentIndex offers a new perspective on cross-layer indexing: sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Cross-Layer Sharing
Indexer Cache
Token Selection
Latent Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Cross-Layer Sharing
Latent Cache
Hierarchical Selection
Multi-head Latent Attention
🔎 Similar Papers