🤖 AI Summary
This study addresses the limitation of existing video generation models in rendering physically plausible interactions and state transitions. To this end, we propose a video generation framework based on a controllable interaction synthesis dataset. Specifically, we first construct a structured interaction taxonomy and leverage image editing models to generate start- and end-state anchors. Subsequently, we introduce State-Guided Sampling to achieve seamless video synthesis. Furthermore, an automated evaluation pipeline aligned with human judgment is designed to optimize data quality. Experimental results demonstrate that fine-tuning base models with our approach yields significant improvements in both the physical plausibility and visual quality of generated interactive videos.
📝 Abstract
While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To address this, we introduce a framework to generate a scalable synthetic dataset of controllable interactions. Our pipeline leverages a structured taxonomy and state-of-the-art image editing models to create explicit `start' and `end' state images, which serve as visual anchors for the interaction. To generate a seamless video utilizing these anchors, we propose State-Guided Sampling (SGS), a novel sampling technique that mitigates artifacts common in naive conditional generation. Furthermore, we develop and validate a new automated evaluation system that aligns with human judgments to ensure data quality. Experiments show that fine-tuning a base model on our dataset significantly enhances its ability to generate plausible interactions.