AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional world models, which optimize solely for factual prediction and struggle to distinguish candidate actions, thereby rendering counterfactual Model Predictive Control (MPC) planning ineffective. To overcome this, we propose AD-WM, which enhances the representation of action discrepancies through action-discriminative joint embeddings. Methodologically, AD-WM incorporates inverse dynamics and normalized reconstruction objectives, combined with residual latent dynamics, conditional mutual information motivation, and a V-JEPA 2 encoder, compelling the latent space to preserve action-relevant information without modifying the test-time MPC pipeline. Experimental results demonstrate that this approach dramatically increases the success rate on the OGBench benchmark from 3.7% to 52%, while achieving a 71.1% zero-shot grasping success rate on real-world robots.
📝 Abstract
Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at https://ad-wm.github.io/.
Problem

Research questions and friction points this paper is trying to address.

World Models
Model Predictive Control
Counterfactual Planning
Action Discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action-Discriminative World Models
Counterfactual Model Predictive Control
Inverse Dynamics Regularization
Conditional Mutual Information
Zero-shot Transfer
🔎 Similar Papers
No similar papers found.