LOCI: Spatial Linear Memory for Streaming World Models
This study addresses the challenges of low scene revisit fidelity and memory scaling with sequence length in video world models by proposing a hybrid spatial memory architecture. The method pioneers the integration of viewpoint-conditioned read-write operations into recurrent linear attention, alternating camera geometric projections with full-attention blocks and incorporating key-value cache management to enable efficient long-range memory retrieval. Experimental results demonstrate that this architecture significantly improves content reproduction fidelity on the MIND benchmark while reducing peak VRAM consumption by approximately 30%. Furthermore, it supports streaming generation of long videos under constant memory constraints.