retrieval-augmented video generation

Designs and implements video synthesis systems that augment generative models with searchable external latent memories or databases and retrieve relevant historical latent representations to condition frame- or clip-level generation. Work includes building dynamic, queryable latent memory structures and retrieval mechanisms, non-local temporal/context conditioning, and techniques to prevent accumulated appearance errors and preserve temporal coherence.

retrieval-augmentedvideogeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation

Dec 21, 2025
TZ
Tianrui Zhu
🏛️ Tsinghua Shenzhen International Graduate School | Tsinghua University

Long-video generation faces dual challenges: poor historical scene consistency and excessive memory consumption—windowed attention causes catastrophic forgetting, while full-history modeling incurs GPU memory bottlenecks. To address this, we propose Memorize-and-Generate (MAG), the first framework to decouple memory compression from frame generation, featuring a lightweight KV cache compression mechanism and a co-training architecture. We introduce MAG-Bench, the first benchmark explicitly designed to evaluate historical memory retention. Additionally, MAG incorporates frame-level autoregressive modeling, enhanced windowed attention, and a novel historical consistency constraint loss. Experiments demonstrate that MAG maintains state-of-the-art (SOTA) performance on standard metrics while improving historical memory retention by 37% and reducing inference latency by 62% compared to full-history attention.

Addresses catastrophic forgetting in long video generationEnhances scene consistency through decoupled memory and generationReduces memory costs while preserving historical context

Existing video editing methods struggle to maintain long-term semantic and structural consistency, primarily due to outdated contextual memory. This work proposes a decoupled multimodal context memory mechanism that constructs separate RGB and depth memory banks to model appearance semantics and geometric structure independently. By incorporating an edit-aware memory update and retrieval strategy, the method enables temporally and viewpoint-consistent video generation. Experimental results demonstrate that the proposed approach significantly outperforms current state-of-the-art techniques after editing, effectively preserving long-range semantic and structural coherence while exhibiting strong robustness.

consistent video generationcontext memorydisentangled representation

This work addresses the limitations of existing video world models, which rely on explicit RGB-space point cloud memory and suffer from high computational overhead and information loss due to pixel-space reconstruction. The authors propose constructing a persistent 3D cache in diffusion latent space, where latent variables are lifted into 3D via depth-guided inverse projection and subsequent view synthesis and querying are performed entirely within the latent space. This approach is the first to maintain full 3D spatial consistency purely in latent space. By integrating latent-space 3D caching, depth-guided inverse projection, latent warping, and geometric priors from diffusion models, the method achieves high-fidelity reconstruction while accelerating end-to-end video generation by 10.57× and reducing memory consumption by 55×. It attains state-of-the-art performance on WorldScore and demonstrates strong results on RealEstate10K.

3D spatial consistencycomputational efficiencylatent representation

This work addresses the challenge of spatial inconsistency during scene revisiting and degraded generation quality in long-term video synthesis, which often arises from the tight coupling between memory modeling and the generative process. To resolve this, the authors propose a decoupled memory-control framework that separates memory modeling from video generation. A lightweight memory branch learns spatial consistency from historical observations and injects relevant memory on demand during generation. Key innovations include a novel on-demand memory mechanism, a camera-aware gating strategy, and a decoupled memory-generation architecture, enhanced by hybrid memory representations and per-frame cross-attention. This approach achieves state-of-the-art visual fidelity and spatial coherence while significantly reducing training costs and data requirements.

long-horizon video generationmemory modelingnovel scene exploration

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Jun 03, 2025
JY
Jiwen Yu
🏛️ The University of Hong Kong | Zhejiang University | Kuaishou Technology

Existing interactive long-video generation methods suffer from insufficient historical context modeling, leading to poor scene consistency. To address this, we propose a novel “context-as-memory” paradigm: historical frames are directly treated as retrievable memory, enabling conditional modeling via frame-wise concatenation—eliminating the need for auxiliary control modules. We further design a lightweight memory retrieval mechanism based on camera field-of-view (FOV) overlap, ensuring information integrity while substantially reducing computational overhead. Our approach is embodied in an end-to-end trainable video diffusion architecture. On interactive long-video generation benchmarks, our method significantly outperforms state-of-the-art approaches, demonstrating strong generalization to unseen open-domain scenes and reducing redundant computation by over 40%.

Enhancing scene-consistent memory in long video generationImproving interactive video generation without external control modulesReducing computational overhead with relevant context retrieval

Latest Papers

What's happening recently
View more

This work addresses the challenge of maintaining long-term memory consistency in video world models under complex camera trajectories. To this end, we propose an attention-based implicit memory retrieval mechanism that enables flexible memory access through viewpoint-aware positional encoding, coupled with a lightweight context compression network for efficient long-sequence processing. Our key innovation lies in the first integration of viewpoint information into positional encoding and its synergistic combination with attention mechanisms for implicit memory modeling. To facilitate training and evaluation of long-horizon video world models, we introduce SceneFly, a large-scale synthetic dataset. Experiments demonstrate that our approach achieves state-of-the-art performance across multiple benchmarks and exhibits strong generalization capabilities in open-domain scenarios.

camera trajectoriesgeneralizationlong-term memory

This work addresses the challenge of achieving real-time dynamic novel view synthesis in multi-view streaming video, which demands both long-term temporal consistency and strict real-time performance. The authors propose an online synthesis framework based on test-time training (TTT) that decouples memory update from application frequency: scene memory is updated only periodically, while each frame efficiently reuses existing memory. To handle dynamic deformations, cross-view attention is introduced. Memory quality is preserved through a dedicated memory loss that enforces internalization of scene structure and a memory caching strategy that mitigates weight drift. The method achieves minute-scale online memory adaptation and real-time rendering in scenes with complex human motion, setting a new state-of-the-art in performance.

dynamic sceneslong-horizon memorynovel view synthesis

Long-form video generation faces significant challenges, including error accumulation, attribute drift, and data scarcity, which hinder identity consistency and dynamic coherence. This work proposes a framework for generating videos of effectively unlimited length by fine-tuning diffusion models on short clips to enable autoregressive segment generation. To ensure both local fidelity and long-term consistency, the approach integrates an inter-segment causal attention mechanism that combines bidirectional and unidirectional attention during long-video training. Additionally, a Truncated Rectified Flow (T-RFlow) is introduced to suppress error propagation, while KV caching is employed to enhance inference efficiency. The method achieves, for the first time, high-fidelity, dynamically natural, and identity-consistent single-shot videos spanning multiple minutes, establishing a new state-of-the-art benchmark in long-form video synthesis.

attribute drifterror accumulationidentity consistency

This work addresses the challenges of historical information loss and audio-visual desynchronization in real-time long-form digital human video generation under limited cache constraints. To this end, the authors propose an anchor-guided persistent memory framework that maintains appearance consistency through fixed visual anchors and compresses audio-visual sequences into dynamic states. A modality-specific residual attention mechanism is introduced to enable efficient joint generation, while decoupling persistent memory from local denoising dependencies to support cross-chunk parallel inference. Furthermore, the method incorporates reference-aware feature modulation and an anchor-preserving causal context distillation strategy. Experimental results demonstrate that the proposed approach significantly improves appearance consistency and audio-visual synchronization in long video generation while accelerating autoregressive inference.

appearance consistencyaudio-video synchronizationdigital human generation

Hot Scholars

ML

Mushui Liu

Zhejiang University
Generative ModelsMulti-modal LearningFew-shot Learning