CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决VLA策略执行误差累积问题,提出CereVLA框架,通过轻量级残差修正与后果预测评估相结合的方法,在不重新训练模型的情况下提高执行效率和准确性。
📝 Abstract
Action-chunked vision-language-action (VLA) policies improve inference efficiency, but limited feedback within committed action chunks can lead to accumulated execution errors. Residual adaptation can correct such deviations without retraining the VLA; however, existing corrections are typically optimized for reference-action consistency without explicitly considering their downstream consequences. To address this limitation, we present Cerebellum-Inspired Consequence-Aware Residual Governance (CereVLA), a unified framework that integrates lightweight residual refinement and predictive consequence evaluation into frozen VLA execution. Corrective actions are first generated by flow-based residual refinement, and their short- and interval-horizon consequences are then evaluated by a recurrent state-space model and a history-aware classifier. Residual corrections predicted to be unfavorable are selectively suppressed by a lightweight governor. Comparisons with state-of-the-art methods on LIBERO-10 and LIBERO-GOAL demonstrate the effectiveness of CereVLA. On SO-101, CereVLA increases task success from 57.5% to 90.0% and reduces mean control steps by 19.6% among successful trials, relative to the frozen SmolVLA baseline.
Problem

Research questions and friction points this paper is trying to address.

Action-chunked VLA
Residual Adaptation
Consequence-aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cerebellum-Inspired
Consequence-Aware
Residual Governance
Flow-based Residual Refinement
Recurrent State-Space Model
🔎 Similar Papers
No similar papers found.