Score
Designs, builds, or analyzes algorithms and end-to-end systems that take multi‑speaker audio/speech recordings and (a) detect speaker turn boundaries and segment the signal into speaker-homogeneous regions, and (b) group, label, or verify those segments by producing speaker embeddings, performing clustering/identification, or running verification/classification. This includes diarization-guided segmentation methods, speaker recognition/labeling and verification pipelines, and the underlying diarization algorithms and embeddings used to assign or verify speaker identities.
To address the challenge of extracting target speaker embeddings from long, overlapping multi-speaker speech, this paper proposes an activity-guided end-to-end embedding extraction method. The core innovation lies in explicitly leveraging speaker activity (SA) as a guidance signal—introducing an activity-aware attention masking mechanism that performs conditional attention pooling directly on raw overlapped speech, thereby replacing the conventional paradigm reliant on pre-segmentation into clean single-speaker utterances. The method jointly models acoustic features and SA annotations using an ECAPA-TDNN backbone to learn robust, activity-conditioned speaker embeddings. Evaluated on speaker verification and speaker diarization tasks, the approach achieves significant performance gains under overlapping conditions, effectively overcoming the inherent limitations of traditional methods constrained to isolated single-speaker segments.
This work addresses the challenging speaker diarization (SAD + diarization) problem in meeting scenarios where multi-channel training data are unavailable and microphone count and geometry are unknown. We propose a generic end-to-end framework that jointly leverages TDOA-driven robust speech segmentation and acoustic embedding clustering for speaker identity attribution. The method requires no microphone prior knowledge and uniformly supports both compact arrays and distributed microphone setups. To handle overlapping speech and dynamic speaker movement, we introduce the first spatial-spectral joint feature representation. Segmentation is guided by TDOA delay estimation, while spectral clustering ensures cross-location speaker ID consistency. Experiments demonstrate that our approach significantly outperforms the single-channel pyannote baseline under both microphone configurations—achieving breakthrough improvements in overlap-aware speech segmentation accuracy and speaker ID stability.
This work addresses the challenges of reproducing and extending DiariZen—an open-source state-of-the-art speaker diarization system—stemming from its cross-library and cross-framework dependencies. We propose the first self-contained, visualizable, and code-aligned modular decomposition of the DiariZen pipeline, structured into seven stages: audio preprocessing, WavLM-Large feature extraction (incorporating structured pruning and layer weighting), Conformer-based backend modeling, powerset classification, VBx clustering, and PLDA scoring. Accompanied by executable scripts and visualization examples, our implementation significantly lowers the barrier to entry for researchers, achieves open-source state-of-the-art performance across multiple benchmarks, and enables fully reproducible experimentation and pedagogical demonstration through comprehensive open-source tutorials.
To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).
Traditional speech separation and speaker diarization typically rely on target-speaker priors or predefined speaker counts, limiting their applicability in open-set scenarios. To address this, we propose an end-to-end joint modeling framework that requires neither speaker registration nor assumptions about the number of speakers. Our method automatically identifies and localizes target speakers via enhanced speaker embedding sampling. A two-stage training strategy coupled with an overlap-aware spectral loss explicitly models overlapping speech structure, thereby improving diarization accuracy and robustness to noise. Evaluated on standard benchmarks, our approach achieves a 71% relative reduction in diarization error rate (DER) and a 69% improvement in corrected word error rate (cpWER) over current state-of-the-art methods. These results significantly advance unsupervised, open-set speech separation and diarization.
This work investigates whether multi-speaker ASR corpora can effectively support speaker diarization. We find that the loosely defined speech segment boundaries in ASR data—often derived from automatic segmentation or transcription alignment—conflict with the strict boundary definitions required by diarization benchmarks, leading to inflated diarization error rates (DER), unreliable evaluation, and poor cross-dataset generalization. To address this, we propose a forced-alignment-based boundary normalization method that refines segment start/end points to conform to diarization conventions, coupled with lightweight post-processing to optimize neural diarization model training. Our approach significantly improves diarization performance: it reduces DER in both streaming and offline settings while simultaneously enhancing joint ASR accuracy. Crucially, this study is the first to systematically demonstrate that ASR segment boundary precision critically governs diarization generalization. The proposed framework establishes a reproducible technical pathway for cross-task data reuse in spoken language understanding.
Speaker diarization addresses the fundamental problem of determining “who spoke when,” yet current approaches exhibit substantial errors—particularly under multi-speaker, multilingual, and heterogeneous acoustic conditions—and these errors propagate to downstream tasks. This study systematically evaluates five state-of-the-art end-to-end models—including PyannoteAI and DiariZen—across four diverse multilingual datasets (English, Chinese, German, Japanese, Spanish), totaling 196.6 hours. We identify two primary error sources for the first time: undetected speech segments and speaker identity confusion under high speaker counts. PyannoteAI achieves the best performance with a 11.2% diarization error rate (DER); DiariZen attains the lowest DER (13.3%) among open-source models, establishing it as the most competitive开源 alternative. The work provides an interpretable bottleneck analysis and empirically grounded boundary validation to guide model optimization.
This work addresses the challenge of robust speaker diarization (SD) in multilingual and code-switched scenarios, where low-resource conditions severely degrade SD performance. We propose a language-agnostic end-to-end SD–ASR–NMT joint pipeline. To enhance SD robustness, we introduce a novel multi-kernel consensus spectral clustering framework that integrates lightweight voice activity detection (VAD), fine-tuned ECAPA-TDNN speaker embeddings, multilingual ASR (Whisper/XLS-R), and neural machine translation, augmented by language identification and rule-based post-processing. To our knowledge, this is the first work to empirically validate the engineering feasibility of full-chain co-optimization of SD–ASR–NMT on real-world multilingual mixed audio. Evaluated on the NCIIPC challenge training set, our system reduces diarization error rate (DER) by 32% over baseline methods, supports Hindi, Tamil, English, and their code-switched combinations, achieves an end-to-end real-time factor <1.8×, and significantly improves cross-lingual generalization and system robustness.
This work addresses the challenge of inconsistent speaker permutation across segments in long-form speech separation, which arises from segment-wise processing. To resolve this without requiring additional training, the authors propose a dynamic clustering approach that maintains a continuously updated reference pool of speaker embeddings. By computing cosine similarities between embeddings from the current segment and those in the reference pool, the method predicts cross-segment speaker alignment and incrementally retains the most representative embeddings, enabling plug-and-play consistency. The approach demonstrates strong robustness under challenging conditions—such as unknown numbers of speakers and prolonged silent intervals—and significantly outperforms existing methods in both dense and sparse long-duration speech scenarios.