MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决自动驾驶中的多模态不确定性问题,提出MM-Future模型,通过生成多个场景-动作假设并使用扩散Transformer共同演化,实现高效的多模态联合建模。
📝 Abstract
Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
multi-mode uncertainty
decision-making
scene evolution
coupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-mode joint world-action modeling
modality-aware diffusion Transformer
MM-Tokens
future-conditioned proposal scorer
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Shuai Liu
NIO
H
Hechangle Gong
Beihang University
H
Hao Jiang
NIO
R
Runlin He
NIO
J
Junxiang Zhan
NIO
Kai Huang
Kai Huang
Sun Yat-sen University
Embedded Systems
S
Sheng Yang
NIO
Shaoqing Ren
Shaoqing Ren
USTC & NIO
Computer VisionDeep LearningAutonomous DrivingArtificial Intelligence