Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a key limitation in existing vision-language-action (VLA) models: their use of fixed execution horizons often leads to outdated actions during critical phases, as replanning timing is decoupled from task progress. To overcome this, the authors propose the Bernoulli Continuation Policy (BCP), a lightweight, plug-in continuation head that models horizon selection as a sequential binary decision—whether to continue executing or replan—enabling adaptive control. BCP incorporates inductive biases through ordinality and temporal prefix sharing, and introduces a replanning efficiency reward to optimize trajectory-level performance without fine-tuning the underlying VLA model. Evaluated on RoboTwin 2.0, BCP improves overall task success by 4.06% and boosts performance on low-success tasks by 11.08%. Real-robot experiments demonstrate a substantial increase in peak task success rate from 44% to 84%, with reduced computational overhead.
📝 Abstract
Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic periodic schedule that is independent of task progress. As a result, when no replanning boundary falls before a critical manipulation stage, it is executed from a stale chunk rather than a freshly replanned one. To address this limitation, we propose Bernoulli-Continuation Policy (BCP), a lightweight, plug-and-play framework for adaptive horizon execution that keeps the base VLA frozen. Given a fixed-length action chunk, its continuation head decomposes execution-horizon selection into a sequence of continue-or-replan decisions, which imposes an ordinal, prefix-sharing inductive bias over candidate horizons rather than treating them as independent classes. Since the optimal horizon for each chunk is not observable, we train this head with reinforcement learning from trajectory-level outcomes and introduce a Replanning-Efficiency Reward that jointly rewards task success and efficient VLA usage, discouraging the policy from collapsing to unnecessarily short horizons. On RoboTwin 2.0 with LingBot-VLA as the base policy, BCP improves the average success rate by +11.08% on 13 low-success tasks and from 89.88% to 93.94% (+4.06%) across all 50 tasks. Although trained only under the Clean setting, BCP generalizes to the Randomized setting, raising the average success rate by +4.06%. It also transfers to a different base policy $π_{0.5}$, achieving a better result on LIBERO (+1.7%) and, notably, on the harder LIBERO-PRO (+6.8%). On a real robot, BCP lifts success from 74% to 92% and from 44% to 84% on two manipulation tasks. Meanwhile, its negligible overhead, combined with higher success, makes BCP's overall runtime even lower than the fixed-horizon baselines.
Problem

Research questions and friction points this paper is trying to address.

adaptive horizon execution
replanning
Vision-Language-Action models
execution horizon
task progress
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive horizon execution
Bernoulli-Continuation Policy
Vision-Language-Action
reinforcement learning
plug-and-play framework
🔎 Similar Papers
No similar papers found.