🤖 AI Summary
This study addresses the challenge of boundary detection in long-video sign language sentence segmentation caused by smooth visual transitions. To this end, it proposes SignShift, a novel framework that introduces a multi-scale inter-frame difference modeling mechanism to precisely capture subtle motion variations for identifying semantic shifts. Additionally, a segment count prediction module and a global-local cue fusion strategy are designed to effectively mitigate over-segmentation and under-segmentation. Experiments demonstrate that SignShift significantly outperforms existing methods on mainstream benchmark datasets, achieving highly accurate sentence-level segmentation without auxiliary information. This work establishes a new paradigm for continuous sign language understanding.
📝 Abstract
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.