🤖 AI Summary
This study addresses the absence of effective evaluation benchmarks for determining when full-duplex speech models should respond, remain silent, or cease speaking in multi-party conversations. To this end, it proposes the first full-duplex interaction benchmark tailored for multi-party shared dialogues, comprising 2,000 multi-person scenarios. The benchmark employs transcription-free continuous audio inputs and distinguishes between explicit and implicit addressing, integrating multimodal large language models such as MiniCPM-o and Moshi for real-time inference to systematically evaluate interactive decision-making capabilities. Experimental results demonstrate that MiniCPM-o 4.5 achieves leading performance across three metrics. Furthermore, the findings reveal that current speech systems struggle to exhibit significant differentiation in response rates between explicit and implicit requests, unlike their text-based counterparts, thereby providing a valuable reference for future research on full-duplex multi-party interactions.
📝 Abstract
Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.