Sparse Planning in Visual World Models via Cost Gradients

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive action search costs of token-based world models over large-scale spatial grids by proposing COSTGRAD, a training-free, goal-directed token selector for planning. For the first time, this method leverages the gradient norm of downstream control objectives to assess token importance, overcoming traditional reliance on prediction accuracy alone and revealing critical design principles for selector-architecture compatibility. By integrating an AdaLN-conditioned predictor with Cross-Entropy Method (CEM) search, experiments demonstrate that at 50% sparsity, planning performance matches or exceeds full-token baselines while achieving a 2.6× per-step speedup. Combined with reduced search iterations, COSTGRAD yields an approximate 5× overall acceleration, substantially lowering computational overhead.
📝 Abstract
Token-based world models enable fine-grained latent planning, but repeatedly processing large spatial token grids makes action search expensive. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at $50\%$ sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured $2.6\times$ wall-clock speedup per environment planning step. Combining token sparsity with reduced CEM search increases this to a $\sim 5\times$ total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This difference tracks action-pathway drift: gradient-selected removal produces less drift than random removal on AdaLN, but more on concat. These results highlight selector-architecture compatibility as a design axis for sparse world-model planning. Project page and demos: https://ycxuyingchen.github.io/costgrad/
Problem

Research questions and friction points this paper is trying to address.

visual world models
latent planning
token sparsity
computational cost
selector-architecture compatibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cost Gradients
Sparse Planning
Visual World Models
Token Selection
Action-Pathway Drift