🤖 AI Summary
This study addresses the limitation of conventional frame-wise camera pose representations, which struggle to capture motion direction and velocity, thereby hindering text alignment and generation performance. To overcome this, we propose a novel representation that decouples camera trajectories into direction and velocity components. Building upon this formulation, we develop CineGEN, a conditional generative model, alongside CineScript, a dataset enriched with high-level metadata, and introduce a robust evaluation protocol. Our approach effectively demonstrates the value of direction-velocity decoupling for encoding directorial intent. Experimental results indicate that the proposed method significantly outperforms existing baselines in both trajectory-text alignment and text-to-trajectory generation tasks, achieving high-quality synthesis of cinematic camera trajectories.
📝 Abstract
Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGEN, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.