🤖 AI Summary
This study addresses the decoupling of human motion and camera trajectory generation tasks, as well as the lack of explicit panning intensity control. To overcome these limitations, this work proposes an asymmetric unified generation framework that independently synthesizes human motion while conditionally driving camera trajectories. Furthermore, it introduces a novel explicit panning intensity control mechanism based on contrastive trajectory pairs, enabling continuous intensity modulation. Experimental evaluations on the PulpMotion dataset demonstrate that the proposed method significantly improves the plausibility of camera distributions and the quality of visual composition. It also achieves precise control over panning intensity, thereby establishing a new paradigm for joint human motion and camera trajectory generation tasks.
📝 Abstract
Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.