🤖 AI Summary
This study addresses the limitations of existing full-duplex speech benchmarks, which exhibit incomplete coverage of interaction behaviors and lack contextual diversity, thereby hindering comprehensive evaluation of real-time interactive capabilities. To this end, we introduce DuplexAct-Bench, a bilingual benchmark that unifies six complementary interaction behaviors—including proactive initiation and active silence—within a single framework while distinguishing between explicit and implicit contexts. Employing streaming speech processing and a multidimensional automated evaluation pipeline, we assess twelve systems. Experimental results reveal significant disparities across systems in both behavioral execution and temporal dynamics, alongside frequent misalignments between semantic quality and behavioral timing. These findings indicate that current models remain incapable of robustly managing real-time engagement decisions.
📝 Abstract
Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content. Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds. Project page: https://alitaxky.icu/DuplexAct-Bench/