train speaker-attributed asr

Design and implement models and training pipelines that jointly transcribe multi‑speaker or overlapping audio and assign each transcribed segment to the correct speaker identity, effectively combining ASR and speaker‑attribution/diarization. This includes selecting architectures and loss functions to balance recognition and diarization objectives, handling overlapping speech and limited real recordings through data augmentation or transfer strategies, and producing speaker‑tagged transcripts at inference.

trainspeaker-attributedasr

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling

Aug 08, 2025
MA
Md Asif Jalal
🏛️ Samsung R&D Institute UK | Centre for Research and Technology Hellas | Language AI R&D Group | Samsung Electronics

Traditional speech separation and speaker diarization typically rely on target-speaker priors or predefined speaker counts, limiting their applicability in open-set scenarios. To address this, we propose an end-to-end joint modeling framework that requires neither speaker registration nor assumptions about the number of speakers. Our method automatically identifies and localizes target speakers via enhanced speaker embedding sampling. A two-stage training strategy coupled with an overlap-aware spectral loss explicitly models overlapping speech structure, thereby improving diarization accuracy and robustness to noise. Evaluated on standard benchmarks, our approach achieves a 71% relative reduction in diarization error rate (DER) and a 69% improvement in corrected word error rate (cpWER) over current state-of-the-art methods. These results significantly advance unsupervised, open-set speech separation and diarization.

Enrollment-free speaker diarization and separationImproved accuracy in overlapped speech framesRobust speaker representation against noise

This work addresses the challenges of training end-to-end large language models for multi-speaker speech recognition under low-resource conditions, where speaker diarization accuracy is often insufficient. To this end, the authors propose a dual-encoder architecture that separately extracts semantic and speaker-specific features, which are then fused via a feature interleaving mechanism before being fed into the large language model. The approach innovatively incorporates a length-aware speaker ID loss and an adaptive ASR loss threshold strategy to jointly optimize speech recognition and speaker diarization. Evaluated on the AliMeeting and Aishell4 datasets, the proposed system achieves relative improvements of 18% and 24% over the baseline, respectively, demonstrating substantial gains in multi-speaker recognition performance in low-resource scenarios.

automatic speech recognitionlarge language modelsmulti-talker speech recognition

Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge

May 28, 2025
SH
Shangkun Huang
🏛️ Beijing Fosafer Information Technology Co., Ltd. | Institute of Forensic Science | The Institute of Linguistics | Chinese Academy of Social Sciences

To address the joint modeling challenge of speaker diarization under overlapping speech and low-SNR ASR in the MISP 2025 Challenge, this work proposes an adaptive hybrid diarization architecture and ASR-aware observation enhancement. First, we introduce a novel overlap-adaptive hybrid diarization framework integrating end-to-end segmentation (WavLM), traditional clustering (AHC/i-vector), and guided source separation (GSS). Second, we design an ASR-aware feature compensation mechanism to overcome GSS performance degradation in noisy conditions. Third, we construct an end-to-end and modularly coordinated SD-ASR cascaded system. Our approach achieves first place in both Track 2 (character error rate: 9.48%) and Track 3 (cpCER: 11.56%), demonstrating state-of-the-art robustness and effectiveness in realistic meeting scenarios with overlapping speech and low signal-to-noise ratios.

ASR-aware method to improve low SNR performanceHybrid diarization for overlapping speech adaptationIntegrated system for real-world meeting scenarios

USED: Universal Speaker Extraction and Diarization

Sep 19, 2023
JA
Junyi Ao
🏛️ The Chinese University of Hong Kong | National University of Singapore | Shanghai Jiao Tong University

To address the inconsistency and scene mismatch arising from the decoupled treatment of speaker extraction and diarization in complex overlapping speech, this paper proposes the first end-to-end jointly optimized framework that unifies frequency-domain speech separation with time-domain speaker activity annotation. The method integrates deep clustering, mask estimation, speaker activity detection, and waveform-level separation modules, supporting variable numbers of speakers and arbitrary overlap ratios. A bidirectional协同 mechanism enables mutual enhancement between extraction and diarization, breaking away from conventional cascaded pipelines. Evaluated on LibriMix, SparseLibriMix, and the real-world telephone conversation dataset CALLHOME, the approach achieves significant improvements over state-of-the-art methods on both tasks—marking the first demonstration of simultaneous gains in extraction quality (e.g., SI-SNRi) and diarization accuracy (e.g., DER).

Audio SeparationSpeaker DiarizationSpeaker Extraction

Latest Papers

What's happening recently
View more

This work addresses the challenge of robust speaker diarization (SD) in multilingual and code-switched scenarios, where low-resource conditions severely degrade SD performance. We propose a language-agnostic end-to-end SD–ASR–NMT joint pipeline. To enhance SD robustness, we introduce a novel multi-kernel consensus spectral clustering framework that integrates lightweight voice activity detection (VAD), fine-tuned ECAPA-TDNN speaker embeddings, multilingual ASR (Whisper/XLS-R), and neural machine translation, augmented by language identification and rule-based post-processing. To our knowledge, this is the first work to empirically validate the engineering feasibility of full-chain co-optimization of SD–ASR–NMT on real-world multilingual mixed audio. Evaluated on the NCIIPC challenge training set, our system reduces diarization error rate (DER) by 32% over baseline methods, supports Hindi, Tamil, English, and their code-switched combinations, achieves an end-to-end real-time factor <1.8×, and significantly improves cross-lingual generalization and system robustness.

Developed a multilingual audio pipeline for speaker identification and diarizationEnhanced speaker diarization in low-resource and code-mixed scenariosIntegrated complementary modules including speech recognition and translation

This work proposes an end-to-end trainable, single-pass alignment approach to address the challenges of aligning and transcribing overlapping speech from multiple speakers. The method introduces, for the first time, the shuffle product and partially ordered finite-state automata (FSA) to directly model (token, speaker) tuples and marginalize over all possible serialized paths at subword, word, and phrase levels. Implemented within the k2/Icefall framework and integrated with Viterbi alignment, the system achieves highly accurate simultaneous alignment and speaker-attributed transcription on synthetic LibriSpeech overlapping speech data, significantly advancing speech recognition performance in multi-speaker scenarios.

alignmentmulti-talker recordingsoverlapped speech

Hot Scholars

HB

Hui Bu

aishell
Speech Recognition、Speech databases and text corpora、Special topics on speech databases and
DK

Dominik Klement

Brno University of Technology
Automatic Speech RecognitionSpeaker DiarizationMachine Learning
AP

Alexander Polok

Brno University of Technology, Faculty of Information Technology
Machine learning
YF

Ying Fang

Westlake University; Zhejiang University
speech recognition
HS

Haoqin Sun

Nankai University
Affective computingSpeech signal processingAudio understanding