Keeping JEPA World Models Plannable When Little of the Frame Moves

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the planning failure in Joint-Embedding Predictive Architecture (JEPA) world models caused by action-insensitive latent representations under minimal-action scenarios. We propose a repair framework adaptable to language goals without retraining. First, we construct the SLIM benchmark and introduce an action-sensitivity probe for theoretical validation. Subsequently, we incorporate an inverse-dynamics auxiliary loss alongside a shared prediction head, effectively enhancing the encoder’s sensitivity to subtle actions. Experimental results demonstrate that our approach significantly increases success rates on the challenging box-pushing task from 0.3% to 35%, while achieving 84% performance on language-goal navigation. These results substantially outperform existing baselines, enabling reliable language-conditioned planning within the latent space.
πŸ“ Abstract
Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.
Problem

Research questions and friction points this paper is trying to address.

latent world models
JEPA
planning
action-sensitivity
language goals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent World Model
Inverse Dynamics Auxiliary Loss
Action-Sensitivity Probe
Language-Goal Planning
SLIM Benchmark
πŸ”Ž Similar Papers