Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the policy mismatch problem caused by runtime safety shielding in Vision-Language-Action (VLA) models by proposing FailBank, a self-evolving framework. To our knowledge, this work is the first to leverage runtime counterfactual corrections as supervisory signals, integrating an outcome-aware admission mechanism with Control Barrier Function (CBF) safety modules to enable continuous policy self-evolution via LoRA fine-tuning. Experiments on the VLA-Arena benchmark demonstrate that FailBank effectively balances task success with safety assurance, significantly improving task completion rates while substantially reducing cumulative costs. The proposed framework consistently outperforms both base policies and conventional runtime shielding approaches, highlighting its efficacy in aligning safety interventions with policy optimization for embodied agents.
📝 Abstract
Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into corrective targets and retains successful uncorrected actions as quiet anchors for guarded LoRA updates. We evaluate FailBank on the VLA-Arena benchmark across two difficulty levels and two VLA backbones. Compared with the base policies, FailBank improves the joint success-cost operating point. Across the two backbones, FailBank improves task success rate by 8.5 and 6.9 percentage points, while reducing policy-induced cumulative cost by 35.6\% and 23.8\%, respectively. Compared with runtime shielding, FailBank raises task success rate by 25.4 and 9.5 percentage points, while maintaining comparable policy-induced cumulative cost. These results show that runtime feedback can serve as persistent policy supervision rather than only as a temporary action constraint.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Runtime Feedback
Policy-Shield Mismatch
Robotic Manipulation
Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Self-Evolution Framework
Runtime Feedback
Control Barrier Function
LoRA Fine-tuning