LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

📅 2025-07-30
📈 Citations: 0
Influential: 0
📄 PDF

career value

220K/year
🤖 AI Summary
To address weak cross-modal feature adaptation, inefficient interaction fusion, and high computational overhead in structural crack multimodal segmentation, this paper proposes LIDAR, a lightweight adaptive perception fusion framework. Methodologically, it introduces a mask-guided dynamic scanning strategy and a dual-domain (spatial–frequency) collaborative fusion mechanism, integrating visual state-space modeling, an adaptive frequency-domain感知器, dual-pooling fusion, and lightweight dynamic-modulated multi-kernel convolution to achieve efficient complementarity between morphological and textural cues. With only 5.35M parameters, LIDAR achieves state-of-the-art performance across three benchmark datasets: on the light-field depth dataset, it attains F1 = 0.8204 and mIoU = 0.8465—demonstrating superior accuracy–efficiency trade-offs.

Technology Category

Application Category

📝 Abstract
Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba.
Problem

Research questions and friction points this paper is trying to address.

Achieving low-cost pixel-level crack segmentation with multimodal data
Lacking adaptive cross-modal feature fusion in current methods
Integrating morphological and textural cues for clear segmentation maps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight Adaptive Cue-Aware Vision Mamba network
Mask-guided Efficient Dynamic Guided Scanning Strategy
Lightweight Dynamically Modulated Multi-Kernel convolution
H
Hui Liu
Tianjin University of Technology, Tianjin, China
C
Chen Jia
Tianjin University of Technology, Tianjin, China
F
Fan Shi
Tianjin University of Technology, Tianjin, China
X
Xu Cheng
Tianjin University of Technology, Tianjin, China
M
Mengfei Shi
Tianjin University of Technology, Tianjin, China
Xia Xie
Xia Xie
Hainan University
BigdataData MiningKnowledage Graph
S
Shengyong Chen
Tianjin University of Technology, Tianjin, China