Dissecting Advantage-Guided Post-Training for Vision-Language-Action Policies

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过控制实验分离并评估了优势引导的视觉-语言-动作策略后训练中的关键设计选择,提出了一种结合时间差优势构建、分组校准和连续优势加权的有效方法。
📝 Abstract
Advantage-guided reinforcement learning provides a practical way to post-train vision-language-action (VLA) policies using limited robot data. However, its performance depends on several coupled choices, including how critic-derived advantages are constructed, calibrated, and used for policy training. Existing recipes often combine these choices into a single end-to-end procedure, making their individual effects difficult to identify. In this work, we dissect advantage-guided VLA post-training through a controlled empirical study that separates these design choices while accounting for their distinct estimands. We develop stage-specific offline evaluation methods to screen alternative choices efficiently, without requiring extensive real-robot policy evaluations for every possible combination. The staged evaluation identifies a modular recipe that combines temporal-difference advantage construction, group-wise calibration, and continuous advantage weighting. Across four real-world bimanual tasks, the resulting recipe improves mean task progress and success over the SFT initialization by 0.42 and 0.63, respectively. Moreover, the proposed evaluation diagnostics show an overall alignment with downstream real-world performance, supporting their use for interpreting empirical outcomes and selecting advantage-guided post-training designs in practice.
Problem

Research questions and friction points this paper is trying to address.

Advantage-guided Reinforcement Learning
Vision-Language-Action Policies
Post-training
Robot Data
Policy Training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Advantage-guided VLA post-training
Stage-specific offline evaluation
Temporal-difference advantage construction
Group-wise calibration
Continuous advantage weighting
🔎 Similar Papers
No similar papers found.