Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference

📅 2026-01-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of achieving both high fidelity and real-time performance in 3D video conferencing under low-bitrate constraints, where conventional 2D compression discards critical geometric details and implicit rendering methods like NeRF incur prohibitive computational costs. To overcome these limitations, the authors propose a lightweight 3D talking-face compression framework that uniquely integrates the FLAME parametric face model with 3D Gaussian Splatting (3DGS). By transmitting only compact facial metadata and leveraging compressed Gaussian attributes alongside optimized MLP weights for efficient reconstruction, the method drastically reduces bandwidth requirements. It achieves superior rate-distortion performance at extremely low bitrates compared to existing approaches, enabling high-quality, real-time 3D facial communication.

Technology Category

Computer Vision: 3D Computer VisionMachine Learning: Learning on the Edge & Model CompressionData Mining & Knowledge Management: Data Compression

Application Category

User Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications.
Problem

Research questions and friction points this paper is trying to address.

3D talking face
low-bitrate compression
high-fidelity rendering
3D video conferencing
lightweight representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D talking face
low-bitrate compression
3D Gaussian Splatting
FLAME parametric model
neural rendering
🔎 Similar Papers
No similar papers found.
J
Jianglong Li
Shanghai Jiao Tong University, Shanghai, China
J
Jun Xu
Shanghai Jiao Tong University, Shanghai, China
B
Bingcong Lu
Shanghai Jiao Tong University, Shanghai, China
Zhengxue Cheng
Zhengxue Cheng
Assistant Researcher, Shanghai Jiao Tong University
Video and Image CodingComputer VisionImage Quality Assessment
H
Hongwei Hu
AntGroup, Shanghai, China
R
Ronghua Wu
AntGroup, Shanghai, China
Li Song
Li Song
Professor of Electronic Engineering, Shanghai Jiao Tong University
Video CodingImage ProcessingComputer Vision