Score
Designing image-formation models and enhancement methods to mitigate turbid, low‑light, and particulate-specific degradations so visual frontends can robustly extract and match features for recognition and real‑time subsea tasks.
Underwater visual enhancement (UVE) and 3D reconstruction face severe challenges due to inherent optical distortions—namely scattering, absorption, and chromatic aberration—while existing literature lacks a systematic survey on their synergistic development. This paper presents the first unified review of UVE and underwater 3D reconstruction, covering classical models, deep learning, Neural Radiance Fields (NeRF), 3D Gaussian Splatting, and physics-guided approaches. We establish a multi-dimensional evaluation framework and conduct quantitative and qualitative comparisons across mainstream benchmarks, revealing performance boundaries and application suitability under diverse degradation conditions. Furthermore, we identify three key future directions: physics-model integration, cross-domain generalization, and real-time robust reconstruction. This work provides both theoretical foundations and practical guidance for advancing underwater visual understanding and geometric reconstruction.
Underwater images suffer severe degradation due to light absorption, scattering, biofouling, and suspended particulates. This work systematically evaluates how image enhancement affects feature matching performance—a critical prerequisite for visual navigation and SLAM. We propose two task-oriented, quantitative metrics—“local matching stability” and “farthest matchable frame”—to establish the first context-aware evaluation framework tailored to downstream vision tasks such as autonomous navigation and SLAM. Through comprehensive feature matching analysis, metric-based assessment, and end-to-end SLAM validation, we expose a significant performance gap between mainstream enhancement methods and real-world task requirements. Experiments demonstrate that our framework more accurately reflects the practical improvement in trajectory estimation and pose robustness conferred by enhancement, outperforming conventional distortion- or perception-based evaluation paradigms.
This work addresses the severe degradation of visual information in turbid underwater environments and the inadequacy of existing synthetic datasets in realistically modeling such degradation, which leads to distorted model evaluation. To bridge this gap, the authors introduce TUB, the first large-scale benchmark dataset of real-world extremely turbid underwater images, comprising 1,320 images annotated with over 16,000 high-confidence instance masks. Furthermore, they propose Phase Congruency-based Degradation (PCD), a novel metric for quantifying turbidity-induced degradation that overcomes the sensitivity of conventional methods to contrast variations. PCD achieves, for the first time, strong correlation with instance segmentation performance and significantly outperforms existing evaluation metrics on both real and synthetic turbid images, thereby establishing a reliable benchmark for underwater vision tasks.
This work addresses the longstanding challenge in underwater image enhancement of simultaneously achieving real-time performance and high color fidelity, as existing methods either rely on computationally complex models that are difficult to deploy or lightweight architectures that underperform under severe degradation. To overcome this trade-off, the authors propose an efficient real-time enhancement framework that innovatively integrates adaptive red-blue channel compensation, multi-branch reparameterized dilated convolutions, and global color correction guided by statistical priors. Remarkably, the model operates with only 3,880 inference parameters, achieving a throughput of 409 FPS. It outperforms state-of-the-art methods across seven evaluation metrics on eight benchmark datasets, yielding a 29.7% improvement in UCIQE scores, and has been successfully deployed on an ROV platform, significantly enhancing downstream visual task performance.
Marine snow—suspended bright speckles—severely degrades feature matching in underwater videos, while the absence of paired ground-truth data hinders existing denoising approaches. To address this, we propose the first self-supervised pseudo-ground-truth generation framework that requires no real ground truth. Leveraging temporal consistency in video sequences and physical priors of underwater light propagation, our end-to-end differentiable pipeline jointly models motion estimation, inter-frame interpolation, noise characterization, and spatiotemporal constraints to synthesize paired “snow-contaminated / snow-free” training samples. This framework overcomes the long-standing limitation imposed by unpaired data scarcity. Evaluated on multiple real underwater video sequences, our method improves SIFT/ORB feature matching success rates by 32.7%, and surpasses state-of-the-art unsupervised methods by 8.2 dB in PSNR and 0.19 in SSIM.
Underwater images suffer severe degradation due to wavelength-dependent light absorption and scattering, yet existing physics-guided methods are hindered by inaccurate estimation of depth and scattering parameters, resulting in poor generalization. To address this, we propose a physics-guided joint training framework featuring the novel Depth-Decoupled Degradation Model (DDM), which explicitly disentangles veiling light, degradation factors, and scene depth. We further design a three-branch subnetwork and a dual-branch UIEConv module to embed underwater imaging physical priors directly into end-to-end optimization. Our method achieves state-of-the-art PSNR/SSIM performance on real-world underwater scenes—including deep-sea environments with artificial illumination—while simultaneously producing high-fidelity depth maps. Notably, it is the first approach to jointly enhance image quality and depth estimation accuracy, thereby enabling physically consistent, dual-task support for underwater 3D perception.
This work addresses the degradation of visual quality in underwater images captured in real oceanic environments, which is primarily caused by depth-dependent forward scattering blur and marine snow artifacts. To tackle this challenge, the authors propose a staged image enhancement framework that explicitly models the depth-dependent forward scattering effect for the first time and extracts realistic marine snow degradation patterns from authentic underwater imagery. These components are leveraged to generate high-fidelity synthetic data for fine-tuning a Joint-ID network, followed by a lightweight contrast enhancement post-processing step. The approach effectively bridges the domain gap between synthetic and real underwater images, yielding significant improvements in UIQM scores and perceptual clarity on a real-world dataset collected off the coast of Korea, thereby enhancing the usability of underwater imagery for robotic operations.
This work addresses the severe degradation of infrared-visible images in marine environments caused by fog and strong reflections, a challenge exacerbated by the absence of end-to-end collaborative frameworks and real-world multimodal marine datasets. To tackle this, the authors propose a multi-task complementary learning framework (MCLF), introduce the first infrared-visible maritime ship dataset (IVMSD) tailored for marine scenes, and develop three key components: a frequency-spatial enhanced complementary (FSEC) module, a semantic-visual consistency attention (SVCA) module, and a cross-modal guided attention mechanism. These innovations jointly optimize image restoration, multimodal fusion, and semantic segmentation. Experimental results demonstrate that the proposed method significantly improves segmentation accuracy and enhances perception robustness under complex marine conditions on the IVMSD benchmark.
This work addresses the limitation of existing underwater image enhancement methods, which primarily optimize for human visual perception and often fail to recover high-frequency details critical for downstream vision tasks such as semantic segmentation and object detection. To bridge this gap, the authors propose DTI-UIE, a task-driven underwater image enhancement framework that jointly optimizes perceptual quality and task performance through a dual-branch architecture and a task-aware attention mechanism. The study introduces TI-UIED, the first task-oriented underwater image enhancement dataset, and designs a task-aware loss function alongside a multi-stage training strategy. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches across multiple downstream tasks, effectively enhancing the robustness and accuracy of machine vision systems in underwater environments.
This work addresses the challenge of robust real-time 3D reconstruction in turbid or low-light underwater environments, where optical cameras alone often fail. To overcome this limitation, the authors propose a hand-eye photometric-acoustic fusion system that, for the first time, enables MASt3R-based real-time dense reconstruction in real-world high-turbidity waters (0.5–12 NTU). The approach leverages MASt3R to extract dense correspondences from optical images and integrates geometric constraints provided by 3D sonar to significantly enhance both robustness and real-time performance. Evaluated in unstructured turbid settings, the method demonstrates superior accuracy and stability compared to existing baselines, establishing a new benchmark for underwater dense reconstruction under adverse visibility conditions.
This work addresses the challenges of underwater image degradation—such as color distortion, low contrast, and poor visibility—caused by light absorption and scattering. Existing methods often suffer from limited generalization due to rigid physical assumptions or insufficient training data. To overcome these limitations, the authors propose a novel enhancement framework that integrates Retinex theory with language-guided semantic priors. The framework features a prior-free illumination estimator, a cross-modal text alignment module, and a semantic-guided restorer, leveraging CLIP-generated textual descriptions to provide high-level semantic guidance. This study pioneers the incorporation of textual semantics into underwater image enhancement, introduces LUIQD-TD—the first large-scale image-text underwater dataset—and designs an Image-Text Semantic Consistency (ITSS) loss. Experiments demonstrate that the method achieves state-of-the-art or comparable performance against 15 leading approaches across four public benchmarks and a newly curated dataset, significantly improving both visual quality and semantic fidelity.