When Does a Spoken Agent Have Enough Evidence to Act? The PACT-SLM Contract Test

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of streaming speech agents to make premature decisions under insufficient evidence, a limitation that conventional metrics fail to evaluate regarding prefix support. To this end, we propose PACT-SLM, a contract testing framework that introduces the first formal definition of "time-to-first-valid-action," decoupling action identity from timing through controlled evaluation. Experiments employ a WavLM Base Plus probe, integrating paired comparisons, noise rendering, and permutation tests for quantitative analysis. Results demonstrate a post-probe onset semantic accuracy of 26.03%, significantly outperforming textual and acoustic baselines. These findings confirm that action identity and timing constitute independent dimensions, establishing a novel evaluation paradigm for streaming speech decision-making behavior.
📝 Abstract
Streaming spoken agents may take an external action before the available speech supports it, yet final-turn scores do not reveal whether each observed prefix supports that action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled evaluation that assigns a first valid action time and measures action identity and timing separately. The primary diagnostic contains 80 paired contrast groups from four held-out semantic families and 1,600 prefix predictions across clean and 15 dB noise renderings. After correcting a mismatch between randomized branch codes and semantic labels, a refitted WavLM Base Plus probe reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval: 22.14%-29.68%), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. It exceeds matched text, scalar-acoustic, and shuffled-representation probes in post-onset label accuracy, but its score is at the 96th percentile of 100 within-prefix label permutations and below the 97.5th-percentile reference (26.73%). Elapsed time is more onset-exact than WavLM Base Plus (36.25% vs. 23.13%) but less accurate about action identity (9.92% vs. 26.03%). These results show that action identity and timing measure distinct aspects of partial-speech decision behavior.
Problem

Research questions and friction points this paper is trying to address.

spoken agent
streaming speech
partial speech action
turn taking
speech language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

PACT-SLM
Spoken Agent
Partial Speech Evaluation
Action Timing
WavLM