When Intent Arrives Late: A Benchmark for Full-Duplex Speech Models under Delayed Intent Revelation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight in existing full-duplex voice model benchmarks, which fail to evaluate safety and response timing when users reveal malicious intent with a delay. To bridge this gap, we construct LateIntent-Bench, simulating delayed-intent scenarios through shared-prefix prompting and controlled pause techniques. We further propose a matched-pair evaluation method alongside an early-response rate metric to decouple variations in safety selectivity from those in responsiveness. Experiments across four mainstream models reveal distinct response patterns, demonstrating that a 1.5-second silence window effectively mitigates increases in harmful engagement. These findings provide empirical evidence to inform the secure design of full-duplex voice interaction systems.
📝 Abstract
Native full-duplex speech models can respond before a user finishes speaking, making the behavior of the model depend on both the final utterance and when intent-defining information arrives. Existing benchmarks primarily evaluate turn-taking mechanics, whereas safety evaluations typically assume that the complete request is observed prior to response generation. We introduce LateIntent-Bench, a matched-pair benchmark for delayed intent revelation. A shared ambiguous prefix precedes either a benign or a harmful continuation, using a controlled pause to delay when the branches become distinguishable. We define the Premature Response Rate (PRR) to measure whether response onset precedes the revelation of intent. Joint evaluation of harmful and benign engagement distinguishes changes in safety selectivity from general losses in responsiveness. Across 3,136 sessions, four native full-duplex models exhibit distinct response-timing patterns under delayed intent. Three models show increased harmful engagement while maintaining benign responsiveness, whereas one loses engagement with both branches. Inserting a 1.5s silence after intent revelation keeps PRR near the no-pause baseline and substantially reduces changes in harmful engagement. These results demonstrate that evaluating fully specified requests alone fails to capture emerging timing patterns, highlighting the necessity of assessing delayed intent revelation in full-duplex models. The code will be publicly released soon.
Problem

Research questions and friction points this paper is trying to address.

full-duplex speech models
delayed intent revelation
safety evaluation
premature response
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Full-Duplex Speech Models
Delayed Intent Revelation
LateIntent-Bench
Premature Response Rate
Safety Evaluation
🔎 Similar Papers
No similar papers found.