IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages

📅 2026-09-25
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing full-duplex voice agent benchmarks support only English, rendering them inadequate for India's multilingual landscape. To bridge this gap, we extend Full-Duplex-Bench to ten Indian languages, constructing a multilingual benchmark comprising 12,350 samples. Methodologically, we propose a language-agnostic VAD-inspired heuristic for evaluating temporal interactions and establish an open-source transcription-translation pipeline enabling automated cross-lingual LLM-based scoring. Experimental results reveal that commercial APIs exhibit consistent performance across languages yet entail inherent trade-offs between response latency and robustness. Conversely, open-source models face complex compromises among backchanneling behavior, generation quality, and latency.
📝 Abstract
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.
Problem

Research questions and friction points this paper is trying to address.

Full-duplex voice agents
Multilingual benchmarking
Indian languages
Conversational events
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Full-Duplex Voice Agents
Multilingual Benchmark
Voice Activity Detection
Cross-lingual Evaluation
Conversational Events Mining
🔎 Similar Papers
No similar papers found.