Relightable 3D Avatar Reconstruction with Semantic-Adaptive Motion-Illumination Responses

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SAMIRA框架,通过语义自适应的运动-光照响应建模方法解决从单目视频重建可重照明3D头像时面部表情和光照效果不准确的问题。
📝 Abstract
Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modeling of both non-rigid facial motion and illumination-dependent appearance. Existing Gaussian avatar methods commonly rely on globally coupled representations, in which Gaussian primitives share a unified motion or illumination response model. Such uniform modeling neglects the distinct motion patterns and material/reflectance properties of different facial semantic regions, thereby limiting fine-grained animation accuracy and reducing relighting plausibility. To address this limitation, we propose SAMIRA, a 3D Gaussian avatar framework for semantic-adaptive motion-illumination response modeling. For motion response modeling, the Semantic-Adaptive Motion Response module rasterizes current-to-reference mesh displacements into a topology-consistent UV space and leverages facial semantics to route displacement features through semantic-specific modulators, predicting localized Gaussian geometric residuals beyond coarse mesh binding. For illumination response modeling, the Semantic-Adaptive Illumination Response module learns compact diffuse and specular response factors for each facial region, allowing Gaussians in different regions to adapt their illumination responses to novel environment lighting. These response factors are incorporated into deferred physically based shading, providing a lightweight approximation of semantic-dependent illumination effects. Extensive experiments on self-reenactment, cross-reenactment, and relighting demonstrate that SAMIRA improves both fine-grained expression reconstruction and relighting realism over existing methods.
Problem

Research questions and friction points this paper is trying to address.

3D Avatar Reconstruction
Monocular Videos
Non-rigid Facial Motion
Illumination-dependent Appearance
Gaussian Primitives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic-Adaptive
Motion-Illumination Responses
Gaussian Avatars
Facial Semantics
Physically Based Shading
🔎 Similar Papers
2024-07-21IEEE Transactions on Pattern Analysis and Machine IntelligenceCitations: 7
2024-07-15IEEE Transactions on Visualization and Computer GraphicsCitations: 2
J
Jiankuo Zhao
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
Xiangyu Zhu
Xiangyu Zhu
Institute for AI Industry Research, Tsinghua University
Reinforcement learning
J
Jijie Li
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
B
Baiqin Wang
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
S
Shukai Chen
ZKTeco Co., Ltd.
Zhen Lei
Zhen Lei
Associate Professor, OSCO Research Chair in Off-site Construction
Offsite ConstructionConstruction Engineering and Management