🤖 AI Summary
This work addresses the challenge of achieving both high fidelity and real-time performance in 3D video conferencing under low-bitrate constraints, where conventional 2D compression discards critical geometric details and implicit rendering methods like NeRF incur prohibitive computational costs. To overcome these limitations, the authors propose a lightweight 3D talking-face compression framework that uniquely integrates the FLAME parametric face model with 3D Gaussian Splatting (3DGS). By transmitting only compact facial metadata and leveraging compressed Gaussian attributes alongside optimized MLP weights for efficient reconstruction, the method drastically reduces bandwidth requirements. It achieves superior rate-distortion performance at extremely low bitrates compared to existing approaches, enabling high-quality, real-time 3D facial communication.
📝 Abstract
The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications.