IntactWorld: Joint World Modeling with Intact Features

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing video generation models, where feature compression induces structural information loss and undermines real-world logical reasoning. To bridge the resulting manifold gap, this work proposes a Joint World Modeling architecture that leverages uncompressed, complete features for prediction. Furthermore, it introduces a novel "full-to-compact" training paradigm that substitutes raw features with CLS tokens to enable efficient single-branch guidance. By integrating intermediate-layer x0 prediction with low-manifold learning techniques, the proposed method substantially reduces computational overhead. Extensive evaluations demonstrate that this approach surpasses the baseline by 2.46 points on the VBench 2.0 benchmark while decreasing spatial memory consumption by 11.4% and reducing inference latency by 43.8%.
📝 Abstract
While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity $v$ within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature $x_0$ at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4\% and cutting inference latency by 43.8\%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.
Problem

Research questions and friction points this paper is trying to address.

World Modeling
Video Generation
Feature Compression
Structural Information Loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint World Modeling
Intact Features
Clean Feature Prediction
Full-to-Compact Training Paradigm
Video Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Boming Tan
University of Science and Technology of China
X
Xiangdong Zhang
Shanghai Jiao Tong University
Y
Yan Xia
University of Science and Technology of China
Q
Qi Zhu
KOKONI 3D, Moxin Technology
Deyi Ji
Deyi Ji
Tencent; USTC Ph.D.
Multimodal LLMComputer VisionNLP
X
Xue Yang
Shanghai Jiao Tong University
Shaofeng Zhang
Shaofeng Zhang
Southern University of Science and Technology
Learn to Optimize