🤖 AI Summary
This study addresses the tendency of streaming speech agents to make premature decisions under insufficient evidence, a limitation that conventional metrics fail to evaluate regarding prefix support. To this end, we propose PACT-SLM, a contract testing framework that introduces the first formal definition of "time-to-first-valid-action," decoupling action identity from timing through controlled evaluation. Experiments employ a WavLM Base Plus probe, integrating paired comparisons, noise rendering, and permutation tests for quantitative analysis. Results demonstrate a post-probe onset semantic accuracy of 26.03%, significantly outperforming textual and acoustic baselines. These findings confirm that action identity and timing constitute independent dimensions, establishing a novel evaluation paradigm for streaming speech decision-making behavior.
📝 Abstract
Streaming spoken agents may take an external action before the available speech supports it, yet final-turn scores do not reveal whether each observed prefix supports that action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled evaluation that assigns a first valid action time and measures action identity and timing separately. The primary diagnostic contains 80 paired contrast groups from four held-out semantic families and 1,600 prefix predictions across clean and 15 dB noise renderings. After correcting a mismatch between randomized branch codes and semantic labels, a refitted WavLM Base Plus probe reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval: 22.14%-29.68%), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. It exceeds matched text, scalar-acoustic, and shuffled-representation probes in post-onset label accuracy, but its score is at the 96th percentile of 100 within-prefix label permutations and below the 97.5th-percentile reference (26.73%). Elapsed time is more onset-exact than WavLM Base Plus (36.25% vs. 23.13%) but less accurate about action identity (9.92% vs. 26.03%). These results show that action identity and timing measure distinct aspects of partial-speech decision behavior.