🤖 AI Summary
This study addresses the absence of reproducible and interpretable reliability evaluation standards for voice agents in critical interactions by proposing the Inquesto scoring protocol. Departing from heterogeneous metric-mixing paradigms, this method quantifies goal achievement rates through explicit failure events and severity levels. It directly measures audio-temporal faults and conducts semantic evaluations using scenario predicates alongside a fixed model judge. Furthermore, the approach integrates audio analysis, tool trajectory tracing, and multi-condition acoustic testing techniques. Experimental validation across 306 calls under 13 configurations demonstrates the necessity of evidence beyond transcribed text. The protocol and implementation code have been publicly released as open source.
📝 Abstract
Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse. Rather than combining heterogeneous metrics, IS defines explicit failure events and severity levels and evaluates the deployed voice pipeline. Timing failures, including talk-over and delayed responses, are measured directly from audio, while semantic and state-dependent failures are evaluated using scenario predicates, tool traces, and a pinned open-model judge. Diagnostic views of behavior, acoustic robustness, identity handling, and speaker groups accompany the score without being combined into it. Inquesto Score v0.1 evaluates 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations of a reference voice-agent system. Our evaluation shows that reliable measurement requires evidence beyond transcripts, explicit treatment of deployment conditions, and validation of the evaluators used to determine outcomes. We release the protocol, reference implementation, and evaluation records.