🤖 AI Summary
This study investigates whether voice agents can reliably translate fluent conversations into the execution of professional workflows. To this end, it introduces the first benchmark designed to evaluate the workflow dependency capabilities of voice agents, constructing a Voice Workbench environment comprising 120 distinct workflows. By incorporating stateful coordination mechanisms, typed tools, and user simulation strategies, the proposed framework systematically assesses task completion under full-duplex interactions. Experimental results reveal that current state-of-the-art voice agents achieve a Pass@1 rate below 25%, with the best Reliable@3 score reaching only 10.8%. These findings demonstrate that stateful coordination constitutes the core bottleneck limiting the reliable execution of complex professional workflows by existing voice agent systems.
📝 Abstract
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.