WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in interactive video world models, where sliding windows discard historical context while full caching saturates GPU memory. To resolve this, we propose an efficient attention architecture that introduces a novel hybrid sparse attention mechanism, integrating linear global attention with head-adaptive sparse attention. Furthermore, we design a hierarchical semantic indexing strategy for KV cache management across multi-level storage, complemented by customized GPU kernels to accelerate computation. This approach effectively balances long-term temporal consistency with low-latency generation. Evaluated on the VBench-Long and InterVBench benchmarks, our method achieves subject consistency scores of 0.9472 and 0.9668, respectively, significantly outperforming existing state-of-the-art approaches.
📝 Abstract
Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primarily rely on sliding-window mechanisms to bound computational complexity. However, this approach inherently sacrifices historical context, undermining the long-range interactive capabilities. Conversely, maintaining a full-history cache remains computationally prohibitive and memory-intensive: the quadratic complexity of attention leads to excessive computational overhead, while the linear growth of the KV cache inevitably leads to GPU memory saturation. To overcome these limitations, we propose WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. First, we introduce Hybrid Sparse Attention (HSA), which integrates linear global attention supplemented with head-adaptive sparse attention. Additionally, we design a Hierarchical KV Cache (HKV) that organizes historical KV pairs into semantically indexed pages across multi-tier memory, enabling fine-grained retrieval and controlled GPU residency. These two designs are supported by tailored kernels to effectively translate their theoretical efficiency into real-world performance. Extensive experiments on VBench-Long and InterVBench demonstrate that WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.
Problem

Research questions and friction points this paper is trying to address.

Interactive Video World Models
Attention Complexity
KV Cache Management
Long-horizon Generation
Autoregressive Diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Video World Models
Hybrid Sparse Attention
Hierarchical KV Cache
Autoregressive Diffusion
Custom Attention Kernels
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zeyu Zhang
DAMO Academy, Alibaba Group
J
Jinyuan Mao
Zhejiang University
D
Dakai An
Hong Kong University of Science and Technology
Wangbo Zhao
Wangbo Zhao
National University of Singapore
Efficient Deep LearningDynamic Neural NetworkMultimodal Model
Hanfeng Lu
Hanfeng Lu
HKUST
mlsyssystems
J
Jiasheng Tang
DAMO Academy, Alibaba Group; Hupan Lab
Yinghao Yu
Yinghao Yu
Engineer, Alibaba
Resource management in containerized clustersGeneration optimizations for distributed systems
W
Wei Wang
Hong Kong University of Science and Technology
Bohan Zhuang
Bohan Zhuang
Zhejiang University
Efficient AIMLSys