Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入U-GROW方法,利用策略不确定性识别关键状态,以提高基于世界模型的视觉-语言-动作策略优化效率。
📝 Abstract
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based reinforcement learning (MBRL) reduces this cost by using a learned world model to generate rollouts for policy optimization. However, it becomes computationally expensive as VLA policies and world models scale. Existing methods typically treat states equally, overlooking substantial differences in their utility for policy improvement. In this paper, we show that policy uncertainty helps identify states with greater potential for policy improvement. The policy exhibits high uncertainty at only a small subset of states, often during decision-sensitive stages where small action differences can alter task outcomes, suggesting that policy improvements at these states could be particularly valuable. Building on these findings, we introduce U-GROW, a lightweight, plug-and-play sampling layer that directs more model rollouts to these informative states. By modifying only the branched-start distribution, U-GROW can be integrated into existing MBRL pipelines without changing the policy optimization objective. Experiments in both simulated and real-world manipulation tasks demonstrate the efficiency and effectiveness of U-GROW, supporting the use of policy uncertainty to guide experience generation.
Problem

Research questions and friction points this paper is trying to address.

Model-based Reinforcement Learning
Policy Uncertainty
Rollouts
Vision-Language-Action Models
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy uncertainty
model-based reinforcement learning (MBRL)
U-GROW
state prioritization
efficient sampling
🔎 Similar Papers
Y
Yifei Sheng
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China; Cirquar Technologies, Nanjing, China
H
Haoxiang Ren
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China; Cirquar Technologies, Nanjing, China
Zhilong Zhang
Zhilong Zhang
Nanjing University
Reinforcement LearningDeep Learning
H
Haonan Wang
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China; Cirquar Technologies, Nanjing, China
R
Runjie Xu
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China; Cirquar Technologies, Nanjing, China
Y
Yihao Sun
Mila – Quebec AI Institute; Université de Montréal
Nan Tang
Nan Tang
National Institute of Biological Sciences, Beijing
stem cell biologyaginglung diseases
Z
Zhichao Wu
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China
Lei Yuan
Lei Yuan
Nanjing University
Machine LearningReinforcement LearningMulti-Agent SystemsEmbodied AI
Haoxin Lin
Haoxin Lin
Nanjing University
Reinforcement LearningRobotics
Y
Yang Yu
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China