VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text-to-video models struggle to generate physically consistent dynamic content due to their reliance on implicit temporal modeling. This work proposes a dual-engine agent framework that, for the first time, leverages executable Blender code as a procedural intermediate representation. In this approach, an encoding agent generates programs describing scene composition and temporal evolution; a simulation engine executes these programs to produce deterministic spatiotemporal drafts, which are then refined by a video generation engine into photorealistic outputs. By decoupling procedural reasoning from high-fidelity rendering, the method significantly enhances controllability, interpretability, and physical consistency. Trained on a newly curated VideoCoCo-3K dataset comprising draft-instruction-target triplets, the model achieves state-of-the-art performance with scores of 0.558 on PhyGenBench and 77.88 on VBench-2.0.
📝 Abstract
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
Problem

Research questions and friction points this paper is trying to address.

physically-consistent video generation
text-to-video
temporal evolution
executable representation
spatiotemporal consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable code
physically-consistent video generation
chain-of-thought
agentic dual-engine
Blender simulation
🔎 Similar Papers
No similar papers found.