Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of inaccurate target localization in vision-language-action (VLA) models during fine-grained manipulation, often caused by unreliable visual representations and attention artifacts. To this end, we propose the AtVLA framework, which introduces learnable register tokens into VLA for the first time to disentangle embodied spatial information from visual attention. We further design an uncertainty-guided local high-resolution recoding mechanism that enables adaptive visual refinement. Through end-to-end embodied training, attention distribution rectification, and action-conditioned rollback, our method significantly enhances spatial perception accuracy with only a 1.4–1.6× increase in computational overhead. Experiments show that AtVLA improves average success rates from 94.2% to 98.4% on LIBERO and SimplerEnv simulation benchmarks and boosts performance on real-world single-view tasks from 46.5% to 69.0%.
📝 Abstract
Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previously documented in generic Vision Transformers, and further show that, in embodied policies, these artifacts are closely associated with spatial perception capabilities acquired during post-training. As the encoder learns task-relevant information such as object location, depth ordering, and local geometry, limited global-token capacity causes part of this information to spill into low-information patch tokens. We introduce AtVLA, a framework that inserts learnable register tokens into the visual encoder. Trained end-to-end using only embodied data and the original action objective, these registers emerge as dedicated carriers of embodied spatial information, while the remaining patch tokens recover clean and spatially faithful attention distributions crucial for precise target localization and fine-grained contact. Clean attention restores reliable localization, but cannot recover geometric details lost in low-resolution observations. AtVLA therefore couples attention rectification with uncertainty-gated local refinement. The action expert samples multiple action chunks and estimates uncertainty from their disagreement; only for uncertain predictions, action-conditioned attention rollout identifies the task-relevant region, which is cropped, re-encoded at high resolution, and appended to the cached prefix for refined action generation. Across LIBERO, SimplerEnv, and a challenging single-view real-world benchmark, AtVLA improves the average LIBERO success rate from 94.2% to 98.4% and real-world success from 46.5% to 69.0%. The cropping is triggered on approximately 30% of replanning steps, resulting in only 1.4-1.6x the total computation of the base model under the representative deployment setting.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Spatial Perception
Attention Artifacts
Robotic Manipulation
Visual Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Adaptive Visual Refinement
Register Tokens
Uncertainty-Gated Attention
Embodied Spatial Perception
Jin Cui
Jin Cui
Principal Engineer
Embedded SystemOS Kernel & DriverHypervisor & VirtualizationComputer uArch modellingFPGA & EDA
Y
Yanbin Hu
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; School of Software Engineering, Xi’an Jiaotong University
X
Xinyue Long
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; School of Software Engineering, Xi’an Jiaotong University
Linkai Li
Linkai Li
Head of Engineering, Orka Inc
Signal ProcessingSpeech EnhancementBiomedical Optics
B
Boran Zhao
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; School of Software Engineering, Xi’an Jiaotong University
Pengju Ren
Pengju Ren
Professor, Xi'an Jiaotong University