🤖 AI Summary
Voice agents exhibit insufficient robustness under acoustic disturbances such as noise and lack end-to-end evaluation mechanisms spanning from signal processing to interaction. This work proposes TRACE, an acoustic robustness evaluation framework designed for task-oriented human-computer interaction. By integrating acoustic stress simulation with a dual-track dialogue scoring system, the method compares original and stressed recordings to quantify task completion rates, erroneous actions, and recovery capabilities, with a focus on optimizing error prevention and recovery mechanisms. Experimental results demonstrate that TRACE effectively guides agent refinement by significantly reducing the incidence of erroneous actions and enhancing user interaction efficiency in adverse acoustic environments.
📝 Abstract
Voice agents must complete users' tasks despite noise, reverberation, and competing speech. Evaluating agents' robustness therefore requires following how acoustic conditions affect the conversation and the actions taken on the user's behalf. This overview examines what existing benchmarks reveal about agents' ability to complete tasks under acoustic stress and where further task-based evaluation is required. We then introduce TRACE, a practical workflow for designing, running, and interpreting evaluations of acoustic robustness in task-oriented human-agent interactions: the same agent attempts a specified task with an original recording and an acoustically stressed copy, and the resulting conversations are scored for task completion, wrong actions, recovery, and user effort. Finally, we explain how results from these evaluations can guide changes to an agent to prevent wrong actions and improve recovery.