🤖 AI Summary
This study addresses the challenge in video generation where explicit trajectories struggle to control the semantic relationship between the camera and moving subjects, proposing a novel task and paradigm for semantic camera motion control. Methodologically, it pioneers the use of semantic labels in place of explicit trajectories, leveraging reference videos and target labels to drive camera behavior. Technically, the approach integrates shared-basis LoRA, motion-conditioned modulation, and a background consistency loss to enable dynamic interactions while preserving source content fidelity. Experimental results demonstrate that the proposed method achieves a semantic motion success rate of 68.6%, significantly outperforming Vista4D (45.3%), while effectively maintaining subject identity consistency.
📝 Abstract
Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.