FViT: A Focal Vision Transformer with Gabor Filter

๐Ÿ“… 2024-02-17
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 3
โœจ Influential: 1
๐Ÿ“„ PDF
๐Ÿค– AI Summary
To address the high computational cost and weak modeling capability for multi-scale and multi-directional features in vision Transformers (ViTs) for dense prediction tasks, this paper proposes the Focal Vision Transformer (FViT). Methodologically, FViT replaces self-attention with learnable Gabor filters (LGFs) to explicitly encode local directional textures; introduces a biologically inspired Focal Vision (BFV) module that emulates neural mechanisms of visual attentionโ€”namely, focal enhancement and surround suppression; and constructs a lightweight pyramid architecture augmented with a multi-path feed-forward network (MPFFN) to improve feature reuse. Experimental results demonstrate that FViT significantly outperforms mainstream ViT variants on semantic segmentation and object detection benchmarks. At comparable accuracy, it reduces computational cost by 30โ€“50%, while exhibiting strong generalization, high efficiency, and excellent scalability.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
Vision transformers have achieved encouraging progress in various computer vision tasks. A common belief is that this is attributed to the competence of self-attention in modeling the global dependencies among feature tokens. Unfortunately, self-attention still faces some challenges in dense prediction tasks, such as the high computational complexity and absence of desirable inductive bias. To address these issues, we revisit the potential benefits of integrating vision transformer with Gabor filter, and propose a Learnable Gabor Filter (LGF) by using convolution. As an alternative to self-attention, we employ LGF to simulate the response of simple cells in the biological visual system to input images, prompting models to focus on discriminative feature representations of targets from various scales and orientations. Additionally, we design a Bionic Focal Vision (BFV) block based on the LGF. This block draws inspiration from neuroscience and introduces a Multi-Path Feed Forward Network (MPFFN) to emulate the working way of biological visual cortex processing information in parallel. Furthermore, we develop a unified and efficient pyramid backbone network family called Focal Vision Transformers (FViTs) by stacking BFV blocks. Experimental results show that FViTs exhibit highly competitive performance in various vision tasks. Especially in terms of computational efficiency and scalability, FViTs show significant advantages compared with other counterparts. Code is available at https://github.com/nkusyl/FViT
Problem

Research questions and friction points this paper is trying to address.

Efficiency Issues
Complexity Challenges
Orientation and Scale Limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gabor Filters
Bio-inspired Visual System
Reduced Computational Complexity
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Nankai University | Tiangong University
Y
Yulong Shi
College of Artificial Intelligence, Nankai University, 300350, Tianjin, China
M
Mingwei Sun
College of Artificial Intelligence, Nankai University, 300350, Tianjin, China
Y
Yongshuai Wang
School of Artificial Intelligence, Tiangong University, 300387, Tianjin, China
R
Rui Wang
H
Hui Sun
Z
Zengqiang Chen
College of Artificial Intelligence, Nankai University, 300350, Tianjin, China; The Key Laboratory of Intelligent Robotics of Tianjin, Nankai University, 300350, Tianjin, China