M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of a fair evaluation benchmark for full-duplex spoken dialogue systems that supports multi-turn, multilingual, and multidomain interactions. To this end, we introduce a new full-duplex dialogue evaluation benchmark covering English and Japanese across both chit-chat and question-answering tasks, along with the first fine-grained evaluation framework designed explicitly for multi-turn interactions. The framework incorporates multiple context conditions—single-turn, user-history-only, and teacher-forced full context—to systematically analyze the impact of dialogue history on model behavior. Experimental results reveal distinct turn-taking characteristics across models, performance disparities across languages and domains, and the nuanced role of context utilization in system effectiveness, thereby establishing a new benchmark and offering deeper insights for full-duplex dialogue research.
📝 Abstract
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.
Problem

Research questions and friction points this paper is trying to address.

full-duplex spoken dialogue
benchmark
multilingual
multidomain
multi-turn conversation
Innovation

Methods, ideas, or system contributions that make the work stand out.

full-duplex spoken dialogue
multilingual benchmark
multi-turn conversation
dialogue context analysis
turn-taking modeling