🤖 AI Summary
This study addresses the lack of physical consistency in video generation by proposing a training-free framework for fluid-object interaction simulation. The method decouples physical reasoning from appearance synthesis, constructing a dual-agent workflow that bridges a physics simulator with a pretrained diffusion model. By integrating generation-time planning, vision-language models, and region-aware latent wrapping with denoising guidance mechanisms, the framework enables plug-and-play, high-fidelity video generation. Experimental results demonstrate that the proposed approach significantly reduces trajectory errors and fluid endpoint errors, substantially outperforming existing baselines. These performance gains are further corroborated by human preference evaluations.
📝 Abstract
We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.