Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决接触丰富操作中的可靠动作评估问题,提出Imagine-RL方法,通过视觉-扭矩潜世界模型预测未来表示,并结合当前证据和预测结果以更好地评估候选动作。
📝 Abstract
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.
Problem

Research questions and friction points this paper is trying to address.

reliable action evaluation
contact-rich manipulation
future visual and contact consequences
noise-space reinforcement learning
frozen VLA policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

action-conditioned visual-torque imagination
visual-torque latent world model (VTLWM)
residual-confidence-guided cross-attention
🔎 Similar Papers
No similar papers found.