Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of recursive prediction error accumulation in visual planning by proposing the SALT model. Its core innovation lies in introducing state-affine properties into latent-space dynamics modeling for the first time, ensuring that the error propagation operator depends solely on the action sequence and thereby eliminating nonlinear error propagation. Furthermore, SALT integrates a joint-embedding world model with a recursive multi-step rollout supervision mechanism to optimize training. Experimental results demonstrate that SALT improves closed-loop success rates by an average of 10% and significantly reduces the failure rate on the OGBench-Cube task from 23.3% to 2.0%, achieving highly reliable visual planning.
📝 Abstract
Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not capture how prediction errors propagate under recursive rollout. We decompose multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions. We show that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making the error propagation operators depend only on the action sequence. Guided by this result, we introduce SALT (State-Affine Latent Transition), an action-conditioned state-affine dynamics model in which the action modulates both the state transformation and the additive update. We train SALT through recursive multi-step rollout supervision, feeding each predicted latent state back into the transition so that training matches how the model is used during planning. Across four visual planning environments, SALT exhibits $1.48$--$2.19\times$ higher one-step prediction error than the matched LeWM baseline, yet improves closed-loop success in every environment by $10.0$ percentage points on average. On OGBench-Cube, the fraction of episodes that fail with a sharp rise in model-predicted cost after execution decreases from $23.3%$ to $2.0%$.
Problem

Research questions and friction points this paper is trying to address.

visual planning
world models
error propagation
multi-step rollout
latent dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

State-Affine Dynamics
Visual Planning
World Models
Multi-step Rollout
Latent Transition