3D-GP-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors

📅 2024-09-06

🏛️ arXiv.org

📈 Citations: 0

✨ Influential: 0

career value

229K/year

🤖 AI Summary

To address inaccurate inter-view correlation modeling caused by large disparities in wide-baseline multi-view image compression, this paper proposes 3D-LMVIC—the first learning-based multi-view image compression framework incorporating 3D Gaussian Splatting as a geometric scene prior. Methodologically, it introduces a learnable disparity estimation network for accurate cross-view disparity prediction, jointly optimized with a lightweight depth-map entropy coder and adaptive view-sequence reordering to suppress geometric redundancy and enhance temporal correlation. Evaluated on multiple benchmarks, 3D-LMVIC significantly outperforms HEVC-MV and state-of-the-art learning-based methods, achieving an average BD-rate reduction of 18.7%. It operates at real-time speed (>30 FPS) with robust compression performance, establishing a new paradigm for efficient and reliable multi-view compression in applications such as VR and autonomous driving.

Technology Category

Application Category

📝 Abstract

Multi-view image compression is vital for 3D-related applications. To effectively model correlations between views, existing methods typically predict disparity between two views on a 2D plane, which works well for small disparities, such as in stereo images, but struggles with larger disparities caused by significant view changes. To address this, we propose a novel approach: learning-based multi-view image coding with 3D Gaussian geometric priors (3D-GP-LMVIC). Our method leverages 3D Gaussian Splatting to derive geometric priors of the 3D scene, enabling more accurate disparity estimation across views within the compression model. Additionally, we introduce a depth map compression model to reduce redundancy in geometric information between views. A multi-view sequence ordering method is also proposed to enhance correlations between adjacent views. Experimental results demonstrate that 3D-GP-LMVIC surpasses both traditional and learning-based methods in performance, while maintaining fast encoding and decoding speed.

Problem

Research questions and friction points this paper is trying to address.

Improves disparity estimation in wide-baseline multi-camera systems

Reduces geometric redundancy across multi-view images

Enhances correlation between adjacent views in multi-view sequences

Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses 3D Gaussian Splatting for geometric priors

Introduces depth map compression for redundancy reduction

Implements multi-view sequence ordering for enhanced correlations

🔎 Similar Papers

LM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors