🤖 AI Summary
This study addresses the bowl-shaped deformation and degraded accuracy in 3D reconstruction caused by refraction when using flat-port underwater cameras. To mitigate this, we propose a physics-based image-space refraction pre-correction method that leverages ray-tracing simulations to characterize refractive distortion and construct a geometric correction model. By mapping images to equivalent pinhole perspective projections, the dominant refraction errors are eliminated prior to reconstruction. The core contribution lies in its downstream-agnostic, plug-and-play nature, enabling seamless integration into existing Structure-from-Motion (SfM) and Visual SLAM pipelines without modifying backend algorithms. Experimental results demonstrate that the proposed approach effectively eliminates reconstruction deformation, improves frame registration rates, and maintains low reprojection errors, thereby significantly enhancing system versatility and adaptability.
📝 Abstract
Consumer-grade cameras in flat-port housings are widely used for underwater exploration and mapping of coral reefs and seafloor habitats due to their low cost and accessibility. However, refraction at flat-port interfaces causes bowl-shaped deformation in reconstructed scenes and camera trajectories, compromising the metric accuracy required for mapping and navigation. To remove the dominant refractive distortion before reconstruction, we introduce a physics-based refraction correction in image space. Our method is downstream-agnostic: the refraction-corrected images can be directly used as input to existing reconstruction and SLAM algorithms. We characterize the refractive distortion through ray-tracing simulations and validate our correction on two real underwater datasets with differing scene structures. Compared with conventional and refractive Structure-from-Motion (SfM), our approach removes reconstruction deformation while registering more frames and maintaining low reprojection error. The correction further generalizes across diverse reconstruction and VSLAM backends, demonstrating its broad applicability to downstream vision pipelines.