EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception
This study addresses the fine-grained perception challenges in high-resolution scenarios, namely visual token redundancy and context loss caused by local cropping. We propose a plug-and-play, lightweight evidence-adaptive framework that pioneers the use of human visual search trajectories to supervise evidence density, thereby guiding region re-reading and dynamic token allocation. Furthermore, a sparse coordinate bridging network is designed to integrate local features with the global scene. The proposed method can be deployed without retraining the host backbone network, ensuring strong compatibility. Experimental results demonstrate that our approach significantly improves average fine-grained accuracy across nine backbone models, outperforming purely global methods under identical token budgets.