🤖 AI Summary
This study addresses the memory explosion problem in video world models caused by the unbounded growth of spatial memory during generation. We propose HexMemory, a video world model that employs a fixed-size hex-planar low-rank tensor representation to achieve constant-memory scene memorization. To ensure high-quality memory updates without per-scene optimization, we design a feedforward write mechanism integrating feature warping, confidence-weighted pooling, and learned residual correction, all while maintaining a constant storage capacity. Experiments demonstrate that our model achieves high-quality long-horizon video generation on benchmarks such as WorldScore, exhibiting strong region revisit consistency with strictly constant feature storage overhead.
📝 Abstract
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory warps the existing planes while preserving their dimensions, then integrates new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latent features from HexMemory to condition subsequent video generation. Because the writer processes only observations from the latest chunk, Honeycomb avoids per-scene optimization and repeated processing of the full generation history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust consistency when revisiting previously observed regions, while maintaining constant feature-storage requirements throughout generation. Code and additional visualizations are available on our https://jackswl.github.io/honeycomb/.