Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of performance degradation and high computational cost in multimodal crack segmentation for industrial facilities under arbitrary modality missingness. The authors propose Compass, a lightweight network featuring Degradation Simulation Distillation (DSD), a Modality-Agnostic Feature-Aware Prototype Transformer (FAPT), Evidential Theory-driven Topology-Preserving Fusion (ETPF), and an uncertainty-gated decoder to achieve robust segmentation regardless of missing modalities. Built upon an efficient Needle RWKV backbone, Compass contains only 2.58 million parameters. It achieves state-of-the-art performance across three datasets, with CrackDepth attaining an F1 score of 0.8216 and mIoU of 0.8434 even when 90% of depth modality data is missing.
📝 Abstract
In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computational cost. Existing methods struggle to address semantic degradation caused by missing modalities. We propose Compass, a lightweight network for robust crack segmentation under arbitrary missing modalities. Compass comprises Degradation Simulation Distillation (DSD), Needle Block, and Evidential Topology-Preserving Fusion (ETPF). DSD constructs a degradation simulation stream that mimics more severe missing conditions and performs reciprocal distillation with the original stream, decoupling complete perception from degradation adaptation. Within DSD, Feature-Aware Prototype Transmitter (FAPT) performs modality agnostic prototype-guided feature completion to maintain semantic integrity under incomplete modality conditions. As a lightweight backbone, Needle injects crack-direction cues into WKV modulation and combines connectivity-aware gating with anisotropic context probing for structure-aware modeling. ETPF fuses multimodal features via Dempster-Shafer evidential combination with uncertainty-gated decoding, preserving crack topology while suppressing unreliable features. Experiments on three datasets demonstrate state-of-the-art (SOTA) performance under diverse missing modality scenarios. Even with 90\% depth modality missing on CrackDepth, Compass achieves F1 of 0.8216 and mIoU of 0.8434 with only 2.58M parameters. The code is available at https://github.com/Karl1109/Compass.
Problem

Research questions and friction points this paper is trying to address.

multimodal crack segmentation
missing modalities
semantic degradation
lightweight network
pixel-level performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Degradation Simulation Distillation
Needle RWKV
Evidential Fusion
Modality Missing Robustness
Lightweight Crack Segmentation
🔎 Similar Papers
No similar papers found.
Hui Liu
Hui Liu
Tianjin University of Technology
Computer Vision
C
Chen Jia
Engineering Research Center of Learning-Based Intelligent System (Ministry of Education), Tianjin University of Technology, Tianjin, China
F
Fan Shi
Engineering Research Center of Learning-Based Intelligent System (Ministry of Education), Tianjin University of Technology, Tianjin, China
X
Xu Cheng
Engineering Research Center of Learning-Based Intelligent System (Ministry of Education), Tianjin University of Technology, Tianjin, China
M
Mianzhao Wang
Engineering Research Center of Learning-Based Intelligent System (Ministry of Education), Tianjin University of Technology, Tianjin, China
S
Shengyong Chen
Engineering Research Center of Learning-Based Intelligent System (Ministry of Education), Tianjin University of Technology, Tianjin, China