🤖 AI Summary
How video representations encode physical information remains unclear. This study addresses this gap by constructing a benchmark comprising 8,000 simulated cases to jointly evaluate cross-modal physical alignment and the recoverability of quantitative information for the first time. We propose a multi-task evaluation framework alongside continual contrastive learning and retrieval-augmented generation methods. Our experiments reveal a performance trade-off induced by contrastive training between physical alignment and information recoverability. Furthermore, we demonstrate that lightweight probes can effectively recover physical information from learned representations, and that incorporating retrieved references significantly enhances the physical fidelity of videos generated by MiniMax-H3.
📝 Abstract
Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.