CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of vision-language-action policies to joint failures by proposing a unified capability-aware adaptation framework. The method integrates temporal encoders, Jacobian grounding, cross-joint attention, and self-supervised physical prediction mechanisms to infer actual joint capabilities without requiring failure labels or privileged information, enabling online recovery through conditioned residual reinforcement learning. Experimental results on the LIBERO benchmark demonstrate that this framework improves the success rate from 24.8% to 59.3%, surpassing baselines by 17.4 percentage points. Furthermore, it achieves zero-shot fault generalization across tasks and actuators, with its effectiveness validated on real-world robots.
📝 Abstract
Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning. CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction. The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers. Across 28 LIBERO tasks, CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator, and additional evaluations characterize transfer to unseen fault families and demonstrate recovery on a physical Franka Panda. https://capable-vla.github.io/
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action policies
Joint fault recovery
Embodiment adaptation
Fault-tolerant control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Capability-Aware Adaptation
Self-Supervised Learning
Residual Reinforcement Learning
Fault Recovery
🔎 Similar Papers
No similar papers found.
M
Mohammad Khoshnazar
University of Bremen, Bremen, Germany
M
Mohammad Dehghani Tezerjani
University of North Texas, Denton, TX, USA
D
Deyuan Qu
Toyota Motor North America, InfoTech Labs, Mountain View, CA, USA
Z
Zhiyuan Gao
University of Bremen, Bremen, Germany
Y
Yanxiang Zhan
University of Bremen, Bremen, Germany
J
Jeroen Schafer
University of Bremen, Bremen, Germany
Andrew Melnik
Andrew Melnik
Bremen University
Digital TwinsRoboticsLLMsFoundation Models
Qing Yang
Qing Yang
Associate Professor of Computer Science, University of North Texas
Connected Autonomous VehicleInternet of ThingsSecurity and Trust
Michael Beetz
Michael Beetz
Intitute for Artificial Intelligence, Computer Science Department, University of Bremen
cognitive roboticsAI-based Roboticsplan-based controlsemantic perceptionknowledge processing for Robots