Score
Design and implement models and training pipelines that jointly transcribe multi‑speaker or overlapping audio and assign each transcribed segment to the correct speaker identity, effectively combining ASR and speaker‑attribution/diarization. This includes selecting architectures and loss functions to balance recognition and diarization objectives, handling overlapping speech and limited real recordings through data augmentation or transfer strategies, and producing speaker‑tagged transcripts at inference.
Monaural multi-speaker automatic speech recognition (ASR) faces persistent challenges including data scarcity, ambiguous speaker attribution, and difficulty in overlapping speech recognition, compounded by the absence of a systematic survey on end-to-end (E2E) approaches. To address this gap, we propose the first unified taxonomy for monaural multi-speaker E2E-ASR, explicitly distinguishing and comparatively analyzing SIMO (single-input, multi-output) and SISO (single-input, single-output) architectural paradigms. We further introduce the novel “speaker-consistency hypothesis stitching” framework to enhance robustness in long-form speech modeling. Through comprehensive cross-benchmark evaluation on LibriCSS, AMI, and other datasets, we characterize performance boundaries and identify fundamental error sources across paradigms. Our analysis distills three critical open challenges—robustness, scalability, and real-time inference—providing both theoretical foundations and actionable technical pathways toward practical multi-speaker ASR systems.
Traditional speech separation and speaker diarization typically rely on target-speaker priors or predefined speaker counts, limiting their applicability in open-set scenarios. To address this, we propose an end-to-end joint modeling framework that requires neither speaker registration nor assumptions about the number of speakers. Our method automatically identifies and localizes target speakers via enhanced speaker embedding sampling. A two-stage training strategy coupled with an overlap-aware spectral loss explicitly models overlapping speech structure, thereby improving diarization accuracy and robustness to noise. Evaluated on standard benchmarks, our approach achieves a 71% relative reduction in diarization error rate (DER) and a 69% improvement in corrected word error rate (cpWER) over current state-of-the-art methods. These results significantly advance unsupervised, open-set speech separation and diarization.
This work addresses the challenges of training end-to-end large language models for multi-speaker speech recognition under low-resource conditions, where speaker diarization accuracy is often insufficient. To this end, the authors propose a dual-encoder architecture that separately extracts semantic and speaker-specific features, which are then fused via a feature interleaving mechanism before being fed into the large language model. The approach innovatively incorporates a length-aware speaker ID loss and an adaptive ASR loss threshold strategy to jointly optimize speech recognition and speaker diarization. Evaluated on the AliMeeting and Aishell4 datasets, the proposed system achieves relative improvements of 18% and 24% over the baseline, respectively, demonstrating substantial gains in multi-speaker recognition performance in low-resource scenarios.
To address the joint modeling challenge of speaker diarization under overlapping speech and low-SNR ASR in the MISP 2025 Challenge, this work proposes an adaptive hybrid diarization architecture and ASR-aware observation enhancement. First, we introduce a novel overlap-adaptive hybrid diarization framework integrating end-to-end segmentation (WavLM), traditional clustering (AHC/i-vector), and guided source separation (GSS). Second, we design an ASR-aware feature compensation mechanism to overcome GSS performance degradation in noisy conditions. Third, we construct an end-to-end and modularly coordinated SD-ASR cascaded system. Our approach achieves first place in both Track 2 (character error rate: 9.48%) and Track 3 (cpCER: 11.56%), demonstrating state-of-the-art robustness and effectiveness in realistic meeting scenarios with overlapping speech and low signal-to-noise ratios.
To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).
This work addresses the challenge of robust speaker diarization (SD) in multilingual and code-switched scenarios, where low-resource conditions severely degrade SD performance. We propose a language-agnostic end-to-end SD–ASR–NMT joint pipeline. To enhance SD robustness, we introduce a novel multi-kernel consensus spectral clustering framework that integrates lightweight voice activity detection (VAD), fine-tuned ECAPA-TDNN speaker embeddings, multilingual ASR (Whisper/XLS-R), and neural machine translation, augmented by language identification and rule-based post-processing. To our knowledge, this is the first work to empirically validate the engineering feasibility of full-chain co-optimization of SD–ASR–NMT on real-world multilingual mixed audio. Evaluated on the NCIIPC challenge training set, our system reduces diarization error rate (DER) by 32% over baseline methods, supports Hindi, Tamil, English, and their code-switched combinations, achieves an end-to-end real-time factor <1.8×, and significantly improves cross-lingual generalization and system robustness.
This work proposes an end-to-end trainable, single-pass alignment approach to address the challenges of aligning and transcribing overlapping speech from multiple speakers. The method introduces, for the first time, the shuffle product and partially ordered finite-state automata (FSA) to directly model (token, speaker) tuples and marginalize over all possible serialized paths at subword, word, and phrase levels. Implemented within the k2/Icefall framework and integrated with Viterbi alignment, the system achieves highly accurate simultaneous alignment and speaker-attributed transcription on synthetic LibriSpeech overlapping speech data, significantly advancing speech recognition performance in multi-speaker scenarios.