Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inefficient historical visual memory retrieval and accumulated 3D reconstruction errors in long-horizon camera-controlled video generation by proposing the GEAR framework. This method leverages geometric information as explicit addresses to route attention toward frame-level visual memories, thereby preventing error accumulation caused by global fusion. Furthermore, it introduces a novel geometry-correspondence attention mechanism with an implicit octree structure, combined with token-level matching and visibility evidence accumulation, to achieve precise historical feature injection while effectively eliminating occlusion interference. Built upon diffusion models, GEAR enables the generation of videos up to one minute in length, achieving state-of-the-art performance in visual quality, camera control accuracy, and revisit consistency.
📝 Abstract
Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insight is that geometry need not explain the scene--it only needs to determine where visual memory should be read from, while attention decides what should be recovered. Based on this insight, we introduce GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as an explicit token-level address for visual memory. Rather than fusing historical observations into a persistent global 3D representation, GEAR retains them as frame latents and uses per-frame geometry only to establish token-level correspondences with target views, thereby avoiding persistent error accumulation from global fusion. Guided by these correspondences, Geometric Correspondence Attention (GCA) selectively injects geometrically matched historical features into noisy target patches during denoising. We further introduce an Invisible Octree to accumulate visibility evidence and reject geometrically plausible but occluded correspondences. Extensive experiments demonstrate that GEAR achieves state-of-the-art visual quality, precise camera control, and revisit consistency, enabling minute-long video generation along challenging trajectories.
Problem

Research questions and friction points this paper is trying to address.

long-horizon video generation
camera control
visual memory
geometric error accumulation
memory access
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Enabled Attention Routing
Token-level Visual Memory
Geometric Correspondence Attention
Invisible Octree
Long-Horizon Video Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zesong Yang
State Key Laboratory of CAD&CG, Zhejiang University
Weikai Chen
Weikai Chen
Principal Research Scientist, Tencent America
3D AIGC3D VisionComputer graphicsVLM
L
Liyuan Cui
State Key Laboratory of CAD&CG, Zhejiang University
Lutao Jiang
Lutao Jiang
PhD Student at HKUST(GZ)
3D VisionComputer VisionGen AIAIGC
R
Runze Zhang
LIGHTSPEED
Y
Yingda Yin
LIGHTSPEED
Xiaoyang Huang
Xiaoyang Huang
Shanghai Jiao Tong University
3D VisionNeural RenderingShape Analysis
K
Kai Yan
LIGHTSPEED
K
Keyang Luo
LIGHTSPEED
W
Wangguandong Zheng
X
Xin Wang
LIGHTSPEED
H
Hujun Bao
State Key Laboratory of CAD&CG, Zhejiang University
Zhaopeng Cui
Zhaopeng Cui
Zhejiang University
Computer VisionRoboticsComputer Graphics