Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that video world models generate visually realistic frames yet violate physical laws, proposing a self-evolving framework that overcomes the traditional limitation of natural language in representing physical knowledge. The framework adopts physics language as a shared, optimizable representation across data, training, and generation, enhancing physical consistency by characterizing entities, causality, and evolution dynamics. It introduces a pioneering atomic assertion evaluation system for physical processes, leveraging an agent feedback loop to iteratively refine instructions and convert model deficiencies into textual queries for retrieving missing scenarios. Experiments demonstrate that this approach significantly improves physical plausibility across four mainstream benchmarks, with our open-source Cosmos3-Nano-based model outperforming the proprietary Veo 3.1.
📝 Abstract
Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.
Problem

Research questions and friction points this paper is trying to address.

Video World Model
Physical Plausibility
Physical Knowledge Representation
Natural Language
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video World Model
Physical Representation
Self-Evolving Language
Agentic Loop
PhysCapBench
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13