Score
Designs and implements systems that automatically locate and label pauses in spoken audio, distinguishing silent pauses from filled pauses (e.g., "uh", "um") and other disfluencies while producing timestamps, durations, and counts. Builds acoustic–prosodic feature extractors and classification models to detect, segment, and categorize pause types and disfluencies for use in downstream analysis or processing.
This work addresses the prevalence of disfluencies—such as filler words, repetitions, and false starts—in automatic speech recognition (ASR) outputs, which degrade text readability and hinder downstream task performance. The authors propose an end-to-end multilingual speech repair approach that integrates token-level sequence labeling, instruction tuning, and a contrastive learning objective to steer large language models toward generating semantically coherent and fluent transcriptions. By jointly optimizing these components, the method overcomes limitations of conventional strategies that rely solely on disfluency detection or data augmentation. Extensive experiments demonstrate substantial improvements over strong baselines in Hindi, Bengali, and Marathi, confirming the model’s effectiveness and scalability across diverse low-resource languages.
This study addresses the fine-grained detection of semantic, respiratory, and hybrid (respiratory-semantic) pauses in post-exercise speech, along with exertion-level classification. We introduce the first multi-type pause annotation dataset for post-exercise speech and a hierarchical cascaded modeling framework. Methodologically, we integrate Wav2Vec2’s hierarchical representations with conventional acoustic features (MFCCs/MFBs), and design both single-model and two-stage cascaded architectures compatible with GRU, CNN-LSTM, AlexNet, and VGG16 for joint multi-task learning. Our key contribution lies in the first systematic annotation and joint recognition of all three pause types in post-exercise speech, enhanced by feature–task co-design to improve generalization. Experiments show pause detection accuracies of 89% (semantic), 55% (respiratory), 86% (hybrid), and 73% overall; exertion-level classification achieves 90.5% accuracy—substantially outperforming prior approaches.
This study addresses the challenge of automatic prosodic labeling to support training prosody-controllable text-to-speech (TTS) systems. We propose a multimodal feature fusion framework that jointly integrates representations from speech self-supervised learning (SSL) models, the Whisper encoder, and phoneme-level language models (PnG BERT and PL-BERT)—the first approach to synergistically model acoustic and linguistic features at the phoneme level. Through feature concatenation and end-to-end joint optimization, our method achieves state-of-the-art prosody prediction performance on Japanese: 89.8% accuracy for accent nucleus detection, 93.2% for pitch contour (high/low tone) classification, and 94.3% for phrase boundary (break index) prediction. The approach provides a high-accuracy, scalable, and fully automated solution for prosodic annotation—particularly valuable for low-resource languages—and advances fine-grained prosody modeling for controllable TTS.
Traditional clinical rating scales for assessing formal thought disorder (FTD) in schizophrenia-spectrum disorders are resource-intensive and difficult to scale. Existing automated speech analysis approaches predominantly rely on single-modality features and fail to jointly model temporal dynamics (e.g., pause patterns) and semantic coherence. To address this, we propose the first systematic multimodal framework integrating ASR-derived pause dynamics—such as pause frequency and duration distribution—with semantic coherence metrics computed via pretrained language models. The framework is designed to be robust across diverse clinical contexts. We employ support vector regression (SVR) for feature fusion and evaluation on the TOPSY dataset yields a correlation coefficient of ρ = 0.649 for FTD severity prediction and an AUC of 83.71% for detecting severe cases—both significantly outperforming unimodal baselines. This work establishes a scalable, objective, and quantitatively grounded methodology for automated assessment of speech disorganization in psychosis.
To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).
This study addresses the poor performance of current automatic speech recognition (ASR) systems on disfluent speech—characterized by hesitations, repetitions, and other non-fluencies—which often leads to information loss or hallucination due to the omission of disfluent elements. The work proposes a novel approach that explicitly incorporates disfluency tags into a pre-trained ASR model and combines them with continual learning to enable incremental adaptation across diverse disfluency distributions while mitigating catastrophic forgetting. Experimental results demonstrate that the method significantly enhances robustness on multiple disfluent speech datasets without compromising overall recognition accuracy. Furthermore, the study uncovers a trade-off between learning disfluency markers and recognition performance and identifies a stable cross-attention head mechanism shared across methods, offering new insights into the internal dynamics of ASR models.
This work addresses the limitations of existing audio segmentation approaches, which heavily rely on textual transcriptions while neglecting intrinsic audio signals, the impact of ASR errors, and transcription-free evaluation protocols. To overcome these issues, the authors propose AudioSeg, a purely audio-based segmentation model, and conduct a systematic comparison among text-based models, acoustic features, AudioSeg, and multimodal large language models (MLLMs). They further introduce a novel transcription-free evaluation framework based on temporal alignment. Experimental results demonstrate that AudioSeg significantly outperforms text-dependent methods, with silent pauses emerging as the strongest acoustic cue. Although MLLMs are constrained by context length limitations, they show promise on short audio segments. This study establishes the first transcription-independent evaluation benchmark for audio segmentation and provides in-depth analysis of the interplay among transcription quality, acoustic properties, and model performance.
This work proposes the first single-pass, end-to-end framework for long-form audio understanding, capable of processing audio up to 60 minutes in a unified manner. Addressing the challenges posed by fragmented context and overlapping speakers in scenarios such as meetings and podcasts, the approach jointly integrates automatic speech recognition, speaker diarization, and timestamp generation. It employs a prompt-based context injection mechanism that enhances the accuracy of domain-specific terminology and homophone disambiguation without requiring explicit language identifiers. Built upon the VibeVoice architecture, the model leverages multi-task joint modeling and supports multilingual and code-switching inputs, significantly outperforming existing systems in complex, long-duration settings by achieving high-fidelity transcription and precise speaker attribution.
This work addresses the significant performance degradation of discrete speech representations—such as semantic or phoneme tokens—in noisy environments, which limits the effectiveness of automatic speech recognition (ASR) systems relying on such representations. To mitigate this issue, the authors propose a front-end enhancement system that directly estimates clean discrete speech tokens from noisy input, trained independently of the ASR back-end. The study presents the first systematic comparison among four enhancement paradigms: waveform-to-waveform, token-to-token, continuous self-supervised learning (SSL) features-to-token, and waveform-to-token. The results demonstrate that the waveform-to-token approach consistently outperforms the others. Experiments on the CHiME-4 dataset show that this method not only surpasses alternative enhancement strategies but also exceeds the performance of ASR systems based on continuous SSL features in most conditions.