Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究针对全双工语音对话模型中的轮流发言评估问题,提出了一种基于说话人意图条件的新评估方法TACT,改进了传统的二元窗口规则。
📝 Abstract
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.
Problem

Research questions and friction points this paper is trying to address.

turn-taking
intent-conditioned
full-duplex spoken dialogue models
Innovation

Methods, ideas, or system contributions that make the work stand out.

intent-conditioned evaluation
TACT benchmark
timing kernels
continuous ranked probability score
full-duplex spoken dialogue
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.