LLVD: LSTM-based Explicit Motion Modeling in Latent Space for Blind Video Denoising

📅 2025-01-10
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of sensor noise modeling and insufficient exploitation of motion-temporal dependencies in blind video RAW denoising, this paper proposes a lightweight end-to-end denoising framework operating in the latent space. The core innovation lies in the first-ever integration of LSTM modules into the encoded feature domain to explicitly model inter-frame motion-temporal dependencies—thereby ensuring both denoising continuity and computational efficiency—while eliminating the need for explicit noise priors to enable blind denoising in real-world scenarios. Experiments demonstrate that our method achieves a 0.3 dB PSNR improvement over state-of-the-art methods in the RAW domain, reduces computational complexity by 59%, and exhibits strong robustness against both synthetic and real-world noise. Notably, it significantly enhances video reconstruction quality under challenging conditions such as low illumination and high ISO settings.

Technology Category

Computer Vision: Motion & TrackingMachine Learning: Deep Generative Models & AutoencodersIntelligent Robots: State Estimation

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGResponsible Web: Machine-in-the-loop, human agency and autonomyUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Video restoration plays a pivotal role in revitalizing degraded video content by rectifying imperfections caused by various degradations introduced during capturing (sensor noise, motion blur, etc.), saving/sharing (compression, resizing, etc.) and editing. This paper introduces a novel algorithm designed for scenarios where noise is introduced during video capture, aiming to enhance the visual quality of videos by reducing unwanted noise artifacts. We propose the Latent space LSTM Video Denoiser (LLVD), an end-to-end blind denoising model. LLVD uniquely combines spatial and temporal feature extraction, employing Long Short Term Memory (LSTM) within the encoded feature domain. This integration of LSTM layers is crucial for maintaining continuity and minimizing flicker in the restored video. Moreover, processing frames in the encoded feature domain significantly reduces computations, resulting in a very lightweight architecture. LLVD's blind nature makes it versatile for real, in-the-wild denoising scenarios where prior information about noise characteristics is not available. Experiments reveal that LLVD demonstrates excellent performance for both synthetic and captured noise. Specifically, LLVD surpasses the current State-Of-The-Art (SOTA) in RAW denoising by 0.3dB, while also achieving a 59% reduction in computational complexity.
Problem

Research questions and friction points this paper is trying to address.

Video Denoising
Visual Clarity
Noise Reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

LSTM technology
noise removal
video frame processing
🔎 Similar Papers
2024-07-11Neural Information Processing SystemsCitations: 0