LOCI: Spatial Linear Memory for Streaming World Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of low scene revisit fidelity and memory scaling with sequence length in video world models by proposing a hybrid spatial memory architecture. The method pioneers the integration of viewpoint-conditioned read-write operations into recurrent linear attention, alternating camera geometric projections with full-attention blocks and incorporating key-value cache management to enable efficient long-range memory retrieval. Experimental results demonstrate that this architecture significantly improves content reproduction fidelity on the MIND benchmark while reducing peak VRAM consumption by approximately 30%. Furthermore, it supports streaming generation of long videos under constant memory constraints.
📝 Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Problem

Research questions and friction points this paper is trying to address.

world models
spatial memory
video streaming
memory efficiency
viewpoint retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming world models
spatial linear memory
hybrid architecture
projective camera geometry
key-value cache
🔎 Similar Papers
No similar papers found.
J
Ji Xia
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence
Tingting Liao
Tingting Liao
PhD of MBZUAI
3D Human Generation
X
Xuezhi Liang
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence
H
Hao Li
Mohamed bin Zayed University of Artificial Intelligence, Pinscreen
G
Guangyi Liu
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence