🤖 AI Summary
This study addresses the limitations of existing video world models, including the lack of geometric constraints, poor multi-view consistency, and difficulties in representing uncontrolled background dynamics. To this end, it proposes a geometry-based multi-agent driving world model. Methodologically, a GeoAdapter module is designed to decouple foreground and background dynamics while establishing an explicit 3D state-sharing mechanism. By integrating a diffusion Transformer with action-guided geometry injection, the model achieves highly consistent generation through unified 3D point map reconstruction and progressive memory updating. Experimental results demonstrate that the proposed approach significantly improves visual fidelity and cross-view consistency, and successfully generalizes to multi-agent, multi-camera scenarios. Furthermore, the MA-CARLA dataset is open-sourced to facilitate research on complex interactive driving environments.
📝 Abstract
Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics. To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues. Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents and multi-camera setting.