Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical issue of state hallucinations in Vision-Language-Action (VLA) models, which frequently lead to failures in robotic manipulation. To mitigate this problem, we propose the SOUL framework, which pioneers the use of sparse autoencoders to precisely localize hallucinatory features. Furthermore, it introduces a sparse feature-based strategy unlearning algorithm designed to selectively eliminate erroneous behavioral knowledge associated with these hallucinations. Our approach significantly reduces hallucination-induced error rates and improves task success rates while effectively preserving the model's original manipulation capabilities. By mitigating such vulnerabilities without compromising foundational skills, this work establishes a novel paradigm for enhancing the robustness and reliability of VLA policies in embodied AI applications.
📝 Abstract
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
State hallucination
Robotic manipulation
Sparse features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
State Hallucination
Sparse Autoencoders
Policy Unlearning
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.