🤖 AI Summary
This study addresses the challenge of generating future 4D dynamic scenes from a single event-RGB frame pair, where long-term temporal history and control priors are inherently lacking. To overcome this limitation, the proposed method leverages event streams as motion priors by introducing an event latent enhancement module and a perceptual dynamics space that decouples depth from optical flow while providing feedback constraints on appearance features. Furthermore, it integrates diffusion models with a multi-scale U-Net architecture for multimodal fusion, thereby enforcing geometric and motion consistency. Experimental evaluations on the VKitti2 and DSEC datasets demonstrate state-of-the-art performance. Notably, this work significantly improves the capacity to predict physically plausible and spatiotemporally consistent 4D scenes under severe high-speed motion blur conditions.
📝 Abstract
We propose UniDynamics, a diffusion-based framework for future 4D dynamic scenes (RGB, depth, and optical flow) generation from a single event-RGB pair, without requiring long histories or control priors as in existing methods, while explicitly modeling future motion fields. The core idea is to leverage event streams to offer an alternative motion prior for single-RGB extrapolation, and to enforce geometric and motion constraints throughout generation via multimodal modeling. Specifically, we design an Event Latent Enhancement (ELE) module to align and enhance event latents into diffusion-injectable conditioning features, providing robust initial motion priors and reliable texture/structure cues. We further introduce a Perceptual Dynamics Space (PDS) embedded in the multi-scale U-Net, which decouples and adaptively interacts depth and flow while continuously feeding back constraints to appearance features, improving geometric-motion consistency for physically plausible and spatiotemporally coherent prediction. Experiments on VKitti2 and DSEC demonstrate state-of-the-art performance, producing high-quality, temporally coherent, and 4D-consistent future predictions, especially under challenging high-speed motion blur.