DiarizationLM: Speaker Diarization Post-Processing with Large Language Models

📅 2024-01-07
🏛️ Interspeech
📈 Citations: 7
Influential: 2
📄 PDF
🤖 AI Summary
To address high word-level speaker error rates and poor transcription readability in speaker diarization, this paper proposes DiarizationLM: the first framework to employ large language models (LLMs) as plug-and-play post-processors that jointly correct raw outputs from ASR and diarization systems via structured, text-based prompting. Without modifying or retraining underlying ASR or diarization modules, DiarizationLM supports task-specific prompt engineering and lightweight fine-tuning. On the Fisher and CALLHOME English datasets, it reduces word-level diarization error rates by 55.5% and 44.9%, respectively, while significantly improving semantic coherence and readability of transcriptions. The core contribution is establishing a zero-coupling paradigm for LLMs in diarization post-processing—enabling efficient, flexible, and scalable multimodal speech understanding without architectural or training dependencies on upstream components.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
In this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system. Various goals can be achieved with the proposed framework, such as improving the readability of the diarized transcript, or reducing the word diarization error rate (WDER). In this framework, the outputs of the automatic speech recognition (ASR) and speaker diarization systems are represented as a compact textual format, which is included in the prompt to an optionally finetuned LLM. The outputs of the LLM can be used as the refined diarization results with the desired enhancement. As a post-processing step, this framework can be easily applied to any off-the-shelf ASR and speaker diarization systems without retraining existing components. Our experiments show that a finetuned PaLM 2-S model can reduce the WDER by rel. 55.5% on the Fisher telephone conversation dataset, and rel. 44.9% on the Callhome English dataset.
Problem

Research questions and friction points this paper is trying to address.

Speaker Recognition
Accuracy Improvement
Error Reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

DiarizationLM
Large-scale Language Model
Speaker Recognition
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Google LLC