🤖 AI Summary
This work addresses the lack of behavioral locality in task vector subtraction within closed-loop robotic control, which often fails to precisely suppress a target skill without degrading other capabilities. The authors propose a Target and Control Audit framework—the first systematic evaluation of task vector subtraction in closed-loop multitask vision-language-action (VLA) policies—and uncover three distinct behavioral patterns: target-control disentanglement, resistance, and global collapse, revealing the inherent fragility of locality in this approach. Through comprehensive analyses—including task vector arithmetic, closed-loop evaluation, cosine similarity, norm matching, retain-aware gradient baselines, and single-skill relearning probes—the study demonstrates that only five out of ten skills in LIBERO-Goal can be effectively suppressed, with an average control retention rate of merely 52%, while all interventions significantly impair unrelated control abilities.
📝 Abstract
Task-vector arithmetic offers a closed-form way to modify a model, yet its behavioral locality remains unclear in closed-loop robot control. We present a target-and-control audit of per-skill task-vector subtraction from multitask vision-language-action (VLA) policies. Across all ten LIBERO-Goal skills, subtraction produces three qualitatively different regimes: target-control separation for five skills, resistance for three, and global collapse for two. On held-out initial states, the five suppressible targets remain at 0% success; however, mean baseline-normalized control retention is only 52%, and each target-suppressing edit materially harms at least one nominally unrelated control. Additional Goal panels show separation across tested policies with continuous-regression, discrete-token, and flow-matching action heads, whereas we observe no clean separation on Spatial and control collapse on the tested Object and Long-horizon panels. Mean task-vector cosine does not account for this variation. A matched-norm control identifies a local sign asymmetry around one Goal anchor, while multi-vector outcomes vary with anchor and scale. Retain-aware gradient baselines provide data-dependent comparators but require removal-time data and optimization; subtraction is data- and gradient-free only at edit time, assuming precomputed expert deltas. Finally, a single-skill relearning probe is consistent with behavioral masking, not certified unlearning. These results characterize task-vector subtraction as a fast but brittle intervention and underscore the need for closed-loop target-and-control evaluation when assessing locality in embodied model editing.