WOVEN: Weaving Visual World Modeling into Multimodal LLMs

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal large models generally lack visual transition capabilities in spatial, embodied, and physical reasoning. This work proposes "visual transition reasoning" as a shared training primitive, constructing corresponding training sources and evaluation benchmarks. We innovatively design a supervision selection strategy based on reasoning operations rather than scene actions, validating it through video-pretrained generative models combined with a systematic fine-tuning scheme. Experiments demonstrate that merely 2,000 samples can replace substantial task-specific data, achieving up to a 27.3 percentage point improvement across 22 external benchmarks. These results confirm the reusability and efficiency of visual world modeling capabilities.
📝 Abstract
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Transition Reasoning
Visual World Modeling
Spatial-Temporal Reasoning
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Transition Reasoning
Multimodal Large Language Models
Visual World Modeling
Transfer Learning
Training Recipe