🤖 AI Summary
This study addresses the challenge of detecting stealthy backdoor attacks in Vision-Language-Action (VLA) models during unseen tasks. To this end, we propose TMT, a runtime detection framework that introduces a novel dual-branch detection mechanism grounded in latent dynamics. This mechanism integrates token manifold analysis with latent transition modeling to precisely identify malicious trigger behaviors. Furthermore, a self-distillation strategy is designed to achieve policy purification without requiring a clean reference model. Extensive evaluations across three categories of VLA backdoor attacks on unseen tasks demonstrate that TMT comprehensively outperforms ten baseline methods, establishing a new state-of-the-art in detection performance.
📝 Abstract
Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears. Detecting such activation is difficult because malicious behavior can comprise individually plausible actions, while unfamiliar tasks introduce legitimate changes in observations and behavior. We introduce TMT, a runtime backdoor detector based on Token Manifold and latent Transition modeling. Trained on benign rollouts, its two branches assess input-token structure and prediction errors in adjacent-layer latent dynamics. A suspicious rollout identified by the token manifold branch, once confirmed through latent deviations, guides transition selection for subsequent monitoring. We further explore policy purification through self-distillation: a frozen copy of the backdoored policy provides benign-input actions to supervise a student on paired benign and triggered observations, without requiring a separate clean reference policy. For evaluation, we adapt traditional backdoor detectors and repurpose anomaly and failure detection methods as VLA backdoor detectors. In a post-hoc comparison with ten baselines, TMT achieves state-of-the-art backdoor detection performance on unseen tasks across three VLA backdoor attacks. Our project page is available at https://zzr42.github.io/tmt/.