Score
Designs and implements annotation schemes, labeling tools, and protocols to mark prosodic features in speech audio — including intonation, stress, rhythm, and accent — and conducts manual or automated prosodic labeling and analysis. Builds and evaluates models and algorithms that predict, model, or otherwise analyze prosody from audio and defines metrics and procedures for prosody evaluation and quality assessment.
This study addresses the challenge of automatic prosodic labeling to support training prosody-controllable text-to-speech (TTS) systems. We propose a multimodal feature fusion framework that jointly integrates representations from speech self-supervised learning (SSL) models, the Whisper encoder, and phoneme-level language models (PnG BERT and PL-BERT)—the first approach to synergistically model acoustic and linguistic features at the phoneme level. Through feature concatenation and end-to-end joint optimization, our method achieves state-of-the-art prosody prediction performance on Japanese: 89.8% accuracy for accent nucleus detection, 93.2% for pitch contour (high/low tone) classification, and 94.3% for phrase boundary (break index) prediction. The approach provides a high-accuracy, scalable, and fully automated solution for prosodic annotation—particularly valuable for low-resource languages—and advances fine-grained prosody modeling for controllable TTS.
This work addresses the insufficient syntactic sensitivity of text-to-speech (TTS) systems in prosodic phrase boundary prediction—particularly for syntactically ambiguous constructions such as garden-path sentences—where models over-rely on punctuation and neglect latent syntactic cues. To rigorously assess TTS models’ syntactic awareness, we introduce a psycholinguistic evaluation paradigm. We propose a punctuation-agnostic fine-tuning strategy that compels models to infer implicit syntactic structure. Our methodology integrates controlled fine-tuning of pretrained TTS models, construction of a syntactically annotated prosodic boundary dataset, development of a human-validated prosody labeling protocol, and design of a contrastive ambiguity analysis framework. Results demonstrate significantly improved syntactic consistency in prosodic boundary placement for complex sentences: fine-tuned models better reflect constituent-level syntactic structure and markedly reduce punctuation dependency. This work provides both a novel methodological framework and empirical evidence for enhancing TTS naturalness and alignment with linguistic structure.
This study addresses the underrepresentation of prosodic phrasing in spontaneous speech synthesis by systematically investigating the impact of manual versus automatic prosodic segmentation on non-autoregressive Brazilian Portuguese speech synthesis (FastSpeech 2). Using an open-source dataset licensed under CC BY-NC-ND 4.0, it presents the first comparative evaluation of these two annotation approaches regarding intonation modeling, pause control, and fluency enhancement. Results show that explicit prosodic segmentation yields modest improvements in intelligibility and acoustic naturalness. Both methods successfully reproduce core accent patterns; however, manual annotation—by preserving greater prosodic variability—significantly outperforms automatic segmentation in nuclear pitch contour fidelity and prosodic diversity. This work provides empirical evidence and methodological guidance for fine-grained prosodic modeling in spontaneous speech synthesis.
To address the labor-intensive and inefficient manual prosodic annotation for Indian languages, this paper introduces SIToBI—the first unified prosodic annotation tool specifically designed for syllable-timed Indian multilingual speech. SIToBI supports phoneme-, syllable-, and word-level time-aligned transcription and automatically generates syllable-level fundamental frequency (F0) contours, pause indices, and relative intensity indices. Methodologically, it integrates signal processing techniques—including YAAPT-based pitch extraction and intensity envelope analysis—with rule-driven segmentation algorithms, innovatively modeling cross-linguistic influences to achieve precise syllable boundary alignment. Experimental evaluation on Tamil, Hindi, and Indian English demonstrates that SIToBI achieves prosodic annotation accuracy approaching human performance, exhibits strong cross-linguistic generalizability, and can be rapidly adapted to other syllable-timed languages.
Current text-to-speech (TTS) systems for audiobooks suffer from inadequate modeling of expressive prosody—specifically pitch, loudness, and speaking rate—limiting naturalness and emotional engagement. Method: We propose a language model–driven prosody prediction framework, trained on 93 meticulously aligned book–audiobook text–speech pairs. This is the first systematic study to model audiobook-specific prosodic patterns, employing multi-task regression with explicit decoupling of prosodic attributes for fine-grained control. Results: On a test set of 24 audiobooks, our pitch predictions outperform commercial TTS in 22 books; loudness predictions better match human narration in 23 books. Large-scale human subjective evaluation confirms statistically significant improvements in reading naturalness (p < 0.01) and user preference. This work establishes a scalable prosody modeling paradigm for expressive TTS and releases a high-quality benchmark dataset for audiobook prosody research.
研究通过信息论视角,利用多模态语言模型量化韵律特征与句法结构间的互信息,证明韵律可减少口语中句法不确定性。
本文提出了一种基于低频幅度调制的卷积神经网络方法,直接从语音信号中提取节奏特征,以解决非母语者语音评估中对韵律评价的问题。
This study addresses the challenge of precisely distinguishing turns, feedback, and pauses in spontaneous dialogue by proposing a semi-automatic annotation pipeline. The method integrates voice activity detection, energy filtering, automatic speech recognition, and contextual post-processing to automatically extract turns and feedback while generating consistent initial annotations to facilitate manual review. Experimental results demonstrate that the pipeline achieves an overall F1 score of 0.621 with a boundary error of approximately 0.15 seconds, exhibiting robustness to variations in listening conditions. By standardizing the dialogue annotation workflow, this work significantly enhances the reproducibility of conversational dynamics analysis.
研究探讨了多模态大语言模型在讽刺检测中是否依赖于语调线索,通过控制实验发现模型误判主要基于高音调和不规则停顿的刻板印象。
为解决多维度语音标注成本高、依赖外部服务等问题,提出SpeechAnnotator框架,利用开源工具和多智能体协作实现高效自动标注。