Dynamic Sparse Attention: Access Patterns and Architecture

📅 2026-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance degradation in dynamic sparse attention mechanisms, where per-token selection of key-value cache entries leads to fragmented working sets and poor cache locality, thereby increasing last-level cache (LLC) misses and reducing decoding throughput. To mitigate this issue, the authors propose a service-oriented cache architecture featuring a lightweight indexer that analyzes key-value access patterns, coupled with an LLC reservation mechanism and a fine-grained, token-level LRU replacement policy. This design effectively alleviates cache fragmentation and improves data locality. Experimental evaluation across multiple open-source backbone models demonstrates that the proposed approach significantly reduces LLC stall misses and enhances online inference throughput for dynamic sparse attention models.

Technology Category

Machine Learning: Hardware-aware MLNatural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Dynamic sparse attention (DSA) reduces the per-token attention bandwidth by restricting computation to a top-k subset of cached key-value (KV) entries, but its token-dependent selection pattern introduces a system-level challenge: the KV working set is fragmented, volatile, and difficult to prefetch, which can translate into poor cache locality and stalled decode throughput. We study these effects by implementing a lightweight indexer for DSA-style selection on multiple open-source backbones and logging per-layer KV indices during autoregressive decoding. Our analysis shows a gap in serving DSA backbones - a potential for a high volume of blocking LL (last level) cache miss events, causing inefficiency; we propose a novel LL cache reservation system to save KV tokens in the LL cache between decode steps, combined with a token-granularity LRU eviction policy, and show on the data we collected how this architecture can benefit serving with DSA implemented on different backbones. Finally, we propose directions for future architectural and algorithmic exploration to improve serving of DSA on modern inference platforms.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Sparse Attention
KV Cache
Cache Locality
Decode Throughput
LL Cache Miss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Sparse Attention
KV Cache Optimization
Last-Level Cache Reservation
Token-Granularity LRU
Autoregressive Decoding