AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing world action models, which lack spatial understanding of driving-relevant regions and thereby constrain trajectory planning performance. To overcome this, we propose AffordDrive3D, a novel framework that introduces driving affordances and geometric awareness into world action modeling for the first time. Leveraging a vision-language model (VLM) as its backbone, the proposed method fuses RGB latent representations to jointly predict future drivable areas, collision risks, and three-dimensional geometric structures, significantly enhancing the model's spatial comprehension capabilities. Experimental evaluations demonstrate that AffordDrive3D achieves state-of-the-art performance on the NAVSIM dataset, yielding 91.3 PDMS and 89.9 EPDMS scores. These results validate the effectiveness of the proposed framework in autonomous driving scenarios, highlighting its potential for improving trajectory planning through enhanced spatial reasoning within world action models.
📝 Abstract
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
world-action model
driving affordance
geometry prediction
trajectory planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Driving Affordance
Geometry Prediction
Vision-Language Model
Trajectory Planning