FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video synthesis methods struggle to simultaneously achieve precise trajectory control and high generation fidelity, particularly when integrating static images with dynamic content. This work proposes a unified trajectory-guided conditional generation framework that decouples foreground representations to disentangle intrinsic motion from global displacement. It introduces a parameter-free, spatially aware implicit injection mechanism and leverages a hybrid curriculum learning strategy combining procedurally simulated and real cinematic data. Without requiring 3D reconstruction or additional adapters, the approach enables high-fidelity, spatiotemporally coherent video synthesis that accurately follows prescribed trajectories. The method significantly outperforms current state-of-the-art techniques in visual quality, temporal consistency, and trajectory adherence, while supporting diverse input sources.
📝 Abstract
Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.
Problem

Research questions and friction points this paper is trying to address.

video compositing
motion control
trajectory guidance
generative modeling
spatial placement
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory-guided generation
unified foreground representation
spatial-aware latent injection
video compositing
motion disentanglement
🔎 Similar Papers