🤖 AI Summary
This study addresses the limitation of existing video generation models, where feature compression induces structural information loss and undermines real-world logical reasoning. To bridge the resulting manifold gap, this work proposes a Joint World Modeling architecture that leverages uncompressed, complete features for prediction. Furthermore, it introduces a novel "full-to-compact" training paradigm that substitutes raw features with CLS tokens to enable efficient single-branch guidance. By integrating intermediate-layer x0 prediction with low-manifold learning techniques, the proposed method substantially reduces computational overhead. Extensive evaluations demonstrate that this approach surpasses the baseline by 2.46 points on the VBench 2.0 benchmark while decreasing spatial memory consumption by 11.4% and reducing inference latency by 43.8%.
📝 Abstract
While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity $v$ within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature $x_0$ at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4\% and cutting inference latency by 43.8\%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.