Kinematics-Centric Continuous Sign Language Retrieval with Gloss-Guided Boundary-Aware Alignment

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the semantic ambiguity caused by appearance features in continuous sign language retrieval by proposing a kinematics-centric framework built upon the 3D SMPL-X motion space. The method leverages gloss-level weak supervision to guide local boundary-aware alignment and incorporates visual knowledge distillation to decompose continuous action sequences, thereby completely decoupling RGB dependencies during inference. Experimental results demonstrate that the proposed framework achieves state-of-the-art bidirectional retrieval performance on the CSL-Daily dataset and exhibits strong competitiveness on PHOENIX-2014T. Overall, this work effectively enables precise and robust continuous sign language retrieval.
📝 Abstract
Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.
Problem

Research questions and friction points this paper is trying to address.

Sign Language Retrieval
Sign Language-Text Alignment
Kinematics Representation
Continuous Sign Language
Motion-Language Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kinematics-centric representation
SMPL-X motion
Gloss-guided alignment
Visual distillation
Continuous sign language retrieval