UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of jointly controlling human motion and camera trajectories in complex scenarios involving multiple people, large movements, occlusions, and dynamic cameras. Existing methods struggle due to heterogeneous visual and geometric control signals and high sensitivity to camera estimation errors. To overcome these limitations, we propose UniMoCa, a novel framework that introduces the Motion-Camera Visual Proxy (MCVP)—a unified, identity-agnostic visual representation encoding both 3D human motion and camera paths within a shared visual space, enabling coherent joint control and editing. By incorporating explicit camera trajectory tokens to enhance geometric rendering, our approach eliminates heterogeneous interfaces and significantly improves temporal consistency, control accuracy, and camera robustness. The accompanying MCVP-Video dataset facilitates training on complex motions and multi-view dynamics, achieving state-of-the-art performance with minimal computational overhead.
📝 Abstract
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.
Problem

Research questions and friction points this paper is trying to address.

human video generation
motion control
camera control
visual proxy
multi-person scenes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion-Camera Visual Proxy
Unified Visual Control
Human Video Generation
Camera Trajectory Representation
Identity-Neutral Rendering
🔎 Similar Papers
No similar papers found.