Honeycomb: Constant-Size Scene Memory Representation for Video World Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory explosion problem in video world models caused by the unbounded growth of spatial memory during generation. We propose HexMemory, a video world model that employs a fixed-size hex-planar low-rank tensor representation to achieve constant-memory scene memorization. To ensure high-quality memory updates without per-scene optimization, we design a feedforward write mechanism integrating feature warping, confidence-weighted pooling, and learned residual correction, all while maintaining a constant storage capacity. Experiments demonstrate that our model achieves high-quality long-horizon video generation on benchmarks such as WorldScore, exhibiting strong region revisit consistency with strictly constant feature storage overhead.
📝 Abstract
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our https://jackswl.github.io/honeycomb/.
Problem

Research questions and friction points this paper is trying to address.

Video World Models
Scene Memory
Long-horizon Video Generation
Storage Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video World Models
Constant-Size Memory
Low-Rank Representation
Scene Memory
Long-Horizon Video Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jack Wei Lun Shi
National University of Singapore
K
Kaichen Zhou
Harvard University
H
Haoyu Chen
Harvard University
Y
Yufeng Weng
National University of Singapore
Keane Ong
Keane Ong
PhD Candidate, National University of Singapore, Massachusetts Institute of Technology
Natural Language ProcessingExplainable AIAffective ComputingFinancial AI
Ruojin Cai
Ruojin Cai
PhD, Cornell University
computer visiondeep learning
Hang Hua
Hang Hua
University of Rochester
Computer VisionNatural Language ProcessingMachine Learning
J
Justin K. W. Yeoh
National University of Singapore
M
Mengyu Wang
Harvard University