🤖 AI Summary
This study addresses the challenges of capturing complex gestures and subtle facial expressions in high-fidelity sign language avatar modeling for the Deaf community. To this end, we construct MVSign, the first multi-view Chinese sign language dataset, and introduce a hybrid SMPL-X parameter annotation pipeline. Methodologically, we design a decoupled avatar representation architecture that disentangles the body, head, and hands, combined with a motion-aware sampling strategy to effectively overcome motion blur and pose diversity. Experimental results demonstrate that our approach achieves high-fidelity visual reconstruction on MVSign, significantly enhancing the quality of hand and facial details. Furthermore, the proposed method exhibits strong generalization capabilities to monocular sign language videos in the wild.
📝 Abstract
In this work, we focus on photorealistic sign avatar modeling, which is crucial for effective communication with the Deaf community and is characterized by complex hand gestures and nuanced facial expressions. To this end, we introduce MVSign, the first multi-view Chinese sign language dataset co-designed with Deaf experts, featuring diverse gestures and rich annotations. For precise SMPL-X annotation, we develop a hybrid fitting pipeline that produces accurate body, hand, and facial parameters and can also be applied to the monocular setting. Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity. Extensive experiments demonstrate that our method achieves high-fidelity visual results on MVSign, particularly in detailed hand and facial regions, and generalizes well to in-the-wild monocular sign language videos. Project page: https://naaapi.github.io/PHOSA.