Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of boundary detection in long-video sign language sentence segmentation caused by smooth visual transitions. To this end, it proposes SignShift, a novel framework that introduces a multi-scale inter-frame difference modeling mechanism to precisely capture subtle motion variations for identifying semantic shifts. Additionally, a segment count prediction module and a global-local cue fusion strategy are designed to effectively mitigate over-segmentation and under-segmentation. Experiments demonstrate that SignShift significantly outperforms existing methods on mainstream benchmark datasets, achieving highly accurate sentence-level segmentation without auxiliary information. This work establishes a new paradigm for continuous sign language understanding.
📝 Abstract
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prerequisite for downstream recognition and translation tasks. However, sentence transitions in sign language are often smooth and visually ambiguous, lacking explicit pauses or posture resets. As a result, static frame representations may fail to capture the subtle temporal changes that indicate sentence boundaries. To address this challenge, we propose \textbf{SignShift}, a difference-aware segmentation framework that explicitly models frame-to-frame feature variation as semantic cues for sentence boundary detection. First, to model the feature variation, we design a Temporal Difference Module, which incorporates full-frame, facial, and hand cues, and employs inter-frame differencing to learn multi-scale temporal variations that capture both fine-grained local kinematics and global semantic transitions. Second, to mitigate over- and under-segmentation issues, we design a Segment Count Prediction module, which predicts the number of sentences to guide boundary selection. Extensive experiments on benchmark datasets demonstrate that SignShift substantially outperforms existing methods, validating its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Sign Language Segmentation
Sentence-level Temporal Segmentation
Visual-only
Semantic Shift
Continuous Sign Language Videos
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sign Language Segmentation
Temporal Difference Module
Difference-Aware Framework
Segment Count Prediction
Vis-SSLS
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bowen Guo
State Key Laboratory for Novel Software Technology, Nanjing University
S
Shiwei Gan
State Key Laboratory for Novel Software Technology, Nanjing University
Yafeng Yin
Yafeng Yin
Nanjing University
Multimodal sensingHuman activity recognitionSign language translationSign language production
X
Xiao Liu
State Key Laboratory for Novel Software Technology, Nanjing University
K
Kuizhuang Liu
State Key Laboratory for Novel Software Technology, Nanjing University
Zhiwei Jiang
Zhiwei Jiang
Nanjing University
Natural Language Processing
L
Lei Xie
State Key Laboratory for Novel Software Technology, Nanjing University