๐ค AI Summary
To address the high computational cost and weak modeling capability for multi-scale and multi-directional features in vision Transformers (ViTs) for dense prediction tasks, this paper proposes the Focal Vision Transformer (FViT). Methodologically, FViT replaces self-attention with learnable Gabor filters (LGFs) to explicitly encode local directional textures; introduces a biologically inspired Focal Vision (BFV) module that emulates neural mechanisms of visual attentionโnamely, focal enhancement and surround suppression; and constructs a lightweight pyramid architecture augmented with a multi-path feed-forward network (MPFFN) to improve feature reuse. Experimental results demonstrate that FViT significantly outperforms mainstream ViT variants on semantic segmentation and object detection benchmarks. At comparable accuracy, it reduces computational cost by 30โ50%, while exhibiting strong generalization, high efficiency, and excellent scalability.
๐ Abstract
Vision transformers have achieved encouraging progress in various computer vision tasks. A common belief is that this is attributed to the competence of self-attention in modeling the global dependencies among feature tokens. Unfortunately, self-attention still faces some challenges in dense prediction tasks, such as the high computational complexity and absence of desirable inductive bias. To address these issues, we revisit the potential benefits of integrating vision transformer with Gabor filter, and propose a Learnable Gabor Filter (LGF) by using convolution. As an alternative to self-attention, we employ LGF to simulate the response of simple cells in the biological visual system to input images, prompting models to focus on discriminative feature representations of targets from various scales and orientations. Additionally, we design a Bionic Focal Vision (BFV) block based on the LGF. This block draws inspiration from neuroscience and introduces a Multi-Path Feed Forward Network (MPFFN) to emulate the working way of biological visual cortex processing information in parallel. Furthermore, we develop a unified and efficient pyramid backbone network family called Focal Vision Transformers (FViTs) by stacking BFV blocks. Experimental results show that FViTs exhibit highly competitive performance in various vision tasks. Especially in terms of computational efficiency and scalability, FViTs show significant advantages compared with other counterparts. Code is available at https://github.com/nkusyl/FViT