$τ$-Multilingual: Benchmarking Voice Agents Across Languages

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing voice agent benchmarks are predominantly limited to English, hindering the evaluation of multilingual full-duplex interaction capabilities. This work extends $\tau$-Voice to five languages—Spanish, Portuguese, Hindi, Korean, and Chinese—to introduce the first multilingual full-duplex voice agent benchmark. Leveraging native-speaker review and automated evaluation techniques, we construct a large-scale test set comprising 4,500 calls, with metrics decoupled into task completion, interaction quality, and generation quality. Experiments reveal language-specific failure modes: performance in Korean and Chinese is substantially lower than in English, while Grok achieves superior task completion but ranks lowest in generation quality. The multilingual data packages and validation tools have been released as open source.
📝 Abstract
English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $τ$-Multilingual, extending $τ$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin systems interrupt more often, and both struggle with tools and entities. Grok leads task completion but scores lowest on generation quality, motivating separate task, interaction, and generation reporting. We release language packs, validated judges, and tools for community-built multilingual voice-agent evaluation.
Problem

Research questions and friction points this paper is trying to address.

multilingual voice agents
benchmarking
cross-lingual evaluation
task completion
failure modes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Benchmark
Voice Agents
Full-duplex Evaluation
Task Completion
Open-source Tools
🔎 Similar Papers