Juno: Taming Predictive Latents for Vision-Language-Action Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the latent-control mismatch in vision-language-action models and the failure of teacher calibration under distribution shifts. We propose a unified framework based on action-conditioned JEPA, featuring a single-JEPA multi-role architecture that serves as an alignment backbone, predictive teacher, and dynamic world model across pretraining, policy learning, and deployment. By introducing a dynamic CLS loss and decoupled inference branches, the framework enables test-time adaptive realignment without expert correction. Experiments demonstrate that our method achieves an average success rate of 68.5% on SimplerEnv, improving to 72.7% after test-time adaptation. Furthermore, it maintains 70%–75% success rates on real robots despite distribution shifts such as background variations.
📝 Abstract
Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from $60.9\%$ to $68.5\%$ over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches $72.7\%$; on a real robot, it retains $70\%$--$75\%$ success under background, height, and object shifts where the base policy collapses to $0\%$.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Joint-embedding predictive architectures
predictive latents
embodiment-specific control
distribution shifts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint-embedding predictive architecture
Vision-Language-Action models
Test-time adaptation
Decoupled reasoning branch
Dynamic CLS loss
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13