Score
Designs and implements image-restoration models and modules that recover color, contrast, and detail of underwater photographs by explicitly modeling and compensating for physics-based degradations such as light absorption, scattering, and backscatter. This work produces medium-aware denoising blocks and physics-prior modules — including pseudo-depth-guided and degradation-aware latent or RGB feature restoration components — to reduce color cast, backscatter artifacts, and boundary ambiguity.
Underwater visual enhancement (UVE) and 3D reconstruction face severe challenges due to inherent optical distortions—namely scattering, absorption, and chromatic aberration—while existing literature lacks a systematic survey on their synergistic development. This paper presents the first unified review of UVE and underwater 3D reconstruction, covering classical models, deep learning, Neural Radiance Fields (NeRF), 3D Gaussian Splatting, and physics-guided approaches. We establish a multi-dimensional evaluation framework and conduct quantitative and qualitative comparisons across mainstream benchmarks, revealing performance boundaries and application suitability under diverse degradation conditions. Furthermore, we identify three key future directions: physics-model integration, cross-domain generalization, and real-time robust reconstruction. This work provides both theoretical foundations and practical guidance for advancing underwater visual understanding and geometric reconstruction.
This work addresses the longstanding challenge in underwater image enhancement of simultaneously achieving real-time performance and high color fidelity, as existing methods either rely on computationally complex models that are difficult to deploy or lightweight architectures that underperform under severe degradation. To overcome this trade-off, the authors propose an efficient real-time enhancement framework that innovatively integrates adaptive red-blue channel compensation, multi-branch reparameterized dilated convolutions, and global color correction guided by statistical priors. Remarkably, the model operates with only 3,880 inference parameters, achieving a throughput of 409 FPS. It outperforms state-of-the-art methods across seven evaluation metrics on eight benchmark datasets, yielding a 29.7% improvement in UCIQE scores, and has been successfully deployed on an ROV platform, significantly enhancing downstream visual task performance.
Underwater images suffer from wavelength-dependent attenuation, scattering, and non-uniform illumination, leading to degradation patterns that vary significantly with water type and depth. Existing methods often exhibit limited generalization under complex lighting conditions. To address this, this work proposes DIVER, an unsupervised domain-invariant enhancement framework that integrates empirical correction with physics-guided modeling. DIVER employs IlluminateNet for adaptive brightness enhancement, spectral equalization filtering, channel-adaptive optical correction, and a physics-constrained network, Hydro-OpticNet, to achieve robust restoration across diverse scenarios—including shallow, deep, turbid waters, and scenes with artificial lighting. Evaluated on eight benchmark datasets, the method consistently achieves state-of-the-art or second-best performance, improving UCIQE by over 9%, reducing GPMAE on SeaThru low-light data by more than 4.9%, and significantly enhancing ORB keypoint repeatability and matching accuracy.
To address severe color distortion and blur in underwater imaging caused by light–water interactions, this paper proposes GuidedHybSensUIR, a prior-guided multi-scale hybrid perception restoration framework. Methodologically, it introduces (1) a novel Color Balance Prior, jointly leveraging strong- and weak-supervision modes to guide feature contextualization and decoding; (2) an integrated architecture combining a Detail Restorer for fine-grained texture reconstruction and a Feature Contextualizer for long-range semantic modeling, balancing low-level fidelity and high-level consistency; and (3) the first comprehensive underwater benchmark covering both paired and unpaired scenarios—comprising six diverse datasets—to systematically evaluate 37 methods. Extensive experiments on six real-world test sets demonstrate state-of-the-art performance, outperforming 14 classical and 23 deep learning-based approaches. The code and benchmark are publicly released to foster standardized evaluation in underwater image restoration.
Underwater images suffer severe degradation due to wavelength-dependent light absorption and scattering, yet existing physics-guided methods are hindered by inaccurate estimation of depth and scattering parameters, resulting in poor generalization. To address this, we propose a physics-guided joint training framework featuring the novel Depth-Decoupled Degradation Model (DDM), which explicitly disentangles veiling light, degradation factors, and scene depth. We further design a three-branch subnetwork and a dual-branch UIEConv module to embed underwater imaging physical priors directly into end-to-end optimization. Our method achieves state-of-the-art PSNR/SSIM performance on real-world underwater scenes—including deep-sea environments with artificial illumination—while simultaneously producing high-fidelity depth maps. Notably, it is the first approach to jointly enhance image quality and depth estimation accuracy, thereby enabling physically consistent, dual-task support for underwater 3D perception.
Underwater images commonly suffer from low contrast, blurriness, and chromatic distortion. Existing methods often couple haze removal and color correction in a single model, neglecting their physical independence and synergistic interaction. This paper proposes WaterFormer, a decoupled Vision Transformer architecture: it employs dedicated dehazing and color restoration blocks to model these two degradation processes separately, and introduces a channel fusion block for dynamic inter-block coordination. A soft reconstruction layer, grounded in the underwater imaging physical model, is incorporated to enhance fidelity. Furthermore, we propose a joint optimization strategy combining chromatic consistency loss and Sobel-based color loss to simultaneously preserve color accuracy and structural details. Extensive experiments on multiple benchmark datasets demonstrate that WaterFormer achieves state-of-the-art performance in PSNR, SSIM, and human perceptual evaluation—significantly improving image contrast, sharpness, and color fidelity.
This study addresses the ill-posed nature and lack of scientific confidence in underwater color restoration by investigating its theoretical solvability boundaries. Through mathematical analysis and uncertainty modeling, we establish ideal conditions under which restoration uncertainty is bounded and converges to zero as camera spatial resolution increases. The research demonstrates that high resolution guarantees asymptotic certainty of restoration results under specific constraints, thereby filling the theoretical gap regarding solvability in this domain. Furthermore, this work effectively bridges the cognitive divide between theoretical derivation and empirical validation, providing a rigorous mathematical foundation and reliability assurance for underwater visual restoration.
This work addresses the challenges of underwater 3D reconstruction and appearance recovery, which are hindered by complex optical effects such as wavelength-dependent attenuation and scattering. While existing NeRF-based methods suffer from slow rendering and color distortion, and 3D Gaussian splatting struggles to model volumetric scattering, this paper presents the first pure 3D Gaussian splatting framework that explicitly integrates local underwater optical properties into Gaussian primitives. By employing a dual-branch optimization strategy, the method enforces underwater photometric consistency while naturally recovering in-air appearance—without requiring auxiliary medium networks. It directly models the underwater optical process within the splatting paradigm for the first time, incorporating depth-guided geometric regularization, perceptual-driven losses, and spectral-spatial physical constraints to jointly preserve geometric fidelity and visual realism. Experiments on both standard and custom datasets demonstrate state-of-the-art performance in novel view synthesis and underwater image restoration, all while maintaining real-time rendering capabilities.
This study addresses the lack of a systematic evaluation framework for underwater image reconstruction, which has hindered comprehensive assessment of methods in terms of accuracy, viewpoint consistency, and robustness to varying water conditions. To bridge this gap, the authors propose the first multidimensional evaluation framework that jointly considers reconstruction fidelity, camera motion consistency, and the impact of water quality, accompanied by a newly curated real-world underwater image dataset for benchmarking. Through extensive experiments comparing traditional physics-based scattering models with emerging vision-language models (VLMs), the results demonstrate that VLMs—despite eschewing explicit physical modeling—consistently outperform conventional approaches across all evaluated metrics, achieving superior reconstruction quality and generalization capability.
Underwater images suffer from severe color casts, low contrast, and detail degradation due to wavelength-dependent light absorption and scattering, significantly impairing downstream vision tasks. To address this, we propose the first conditional diffusion-based framework for underwater image enhancement. Our method introduces two key innovations: (1) a chrominance-prior-guided color compensation strategy grounded in optical physics for accurate, interpretable color correction; and (2) a cross-domain consistency loss that jointly optimizes pixel-level fidelity, perceptual quality, structural preservation, and frequency-domain feature alignment. The architecture integrates cross-attention mechanisms, residual dense blocks, and multi-resolution attention to jointly capture global semantics and fine-grained local details. Extensive experiments on multiple benchmark datasets demonstrate that our approach consistently outperforms state-of-the-art CNN-, GAN-, and diffusion-based methods, achieving superior performance in both color correction accuracy and comprehensive image quality metrics.
This work addresses the challenges of underwater image degradation—such as color distortion, low contrast, and poor visibility—caused by light absorption and scattering. Existing methods often suffer from limited generalization due to rigid physical assumptions or insufficient training data. To overcome these limitations, the authors propose a novel enhancement framework that integrates Retinex theory with language-guided semantic priors. The framework features a prior-free illumination estimator, a cross-modal text alignment module, and a semantic-guided restorer, leveraging CLIP-generated textual descriptions to provide high-level semantic guidance. This study pioneers the incorporation of textual semantics into underwater image enhancement, introduces LUIQD-TD—the first large-scale image-text underwater dataset—and designs an Image-Text Semantic Consistency (ITSS) loss. Experiments demonstrate that the method achieves state-of-the-art or comparable performance against 15 leading approaches across four public benchmarks and a newly curated dataset, significantly improving both visual quality and semantic fidelity.