S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出S3VD框架,通过多尺度语义融合和时空扫描融合模块解决雨天视频中高频率细节损坏及运动模糊问题,提高视觉任务可靠性。
📝 Abstract
Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear complexity and ability to model long-range dependencies. However, when confronted with the poor visual representations in rainy videos, Mamba still faces difficulties in preserving the integrity of 2D spatial semantics and modeling 3D spatio-temporal correlations. To break these limitations, we introduce S3VD, a Semantic-Guidance Spatio-Temporal Scanning framework for video deraining, featuring two key innovations: Multi-Scale Semantic Fusion (MSSF) Module and Spatio-Temporal Scanning Fusion (STSF) Module. The former integrates temporal semantic priors from DINOv2 to guide precise feature representation and counteract the loss of local semantic context inherent to Mamba's 1D flatten operation, enhancing robustness against extreme degradation. The latter introduces a spatio-temporal scanning mechanism and devises a Decoupled-Gating Mamba (DG-Mamba) layer, which employs two independent gates to adaptively control preceding and subsequent contextual information within the input clip, optimizing intra-frame and inter-frame correlation modeling. Experiments on video deraining benchmarks demonstrate the superiority of S3VD, achieving state-of-the-art performance with an average 0.84 dB PSNR improvement over Mamba-based baselines.
Problem

Research questions and friction points this paper is trying to address.

heavy rainfall
visual degradation
spatio-temporal correlations
semantic integrity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic-Guidance
Spatio-Temporal Scanning
Multi-Scale Semantic Fusion
Decoupled-Gating Mamba
🔎 Similar Papers
2024-03-05IEEE transactions on circuits and systems for video technology (Print)Citations: 0
Kui Jiang
Kui Jiang
Harbin Institute of Technology
computer visionimage processingdeep learning
Y
Yi’ang Chen
School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China
Y
Yan Luo
School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China
Z
Zhaocheng Yu
School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China
Junjun Jiang
Junjun Jiang
Harbin Institute of Technology
Image ProcessingComputer VisionMachine Learning
X
Xianming Liu
School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China