🤖 AI Summary
This work addresses the absence of speech benchmarks that comprehensively cover all 22 official languages of India and reflect real-world multilingual scenarios for joint speaker diarization and automatic speech recognition (ASR). To bridge this gap, we introduce and publicly release Indic DiarBench, a benchmark dataset comprising 108 hours of naturally occurring multi-speaker audio spanning near-field meetings, far-field recordings, and in-the-wild settings. It is the first dataset to fully encompass all Indian official languages and incorporates complex linguistic phenomena such as code-switching, dialectal variation, and speaker overlap. The dataset includes human-verified, time-aligned transcripts with speaker labels. We further establish baseline systems leveraging commercial ASR APIs and multimodal large language models, providing a standardized evaluation platform to advance research in multilingual joint diarization and ASR and foster more inclusive speech technologies.
📝 Abstract
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.