🤖 AI Summary
This study addresses the disconnection between sentence stress and word stress detection tasks, where shared prosodic cues are often overlooked. To bridge this gap, we propose a joint sentence stress detection framework that incorporates word stress as auxiliary modeling. Methodologically, a word-span stress regularizer is introduced to optimize the loss function by reinforcing token-level probability constraints, thereby enabling effective multi-task collaborative learning. Evaluated on the TinyStress-15K benchmark, the proposed paradigm surpasses existing strong baselines and achieves state-of-the-art performance in sentence stress detection. Overall, this work presents an efficient joint modeling paradigm for speech prosody analysis, demonstrating that leveraging word-level stress information can significantly enhance sentence-level prosodic prediction through unified optimization.
📝 Abstract
Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.