🤖 AI Summary
This study addresses the problem of LLM agent task failures caused by speech transcription errors, noting that conventional ASR metrics fail to capture information preservation. To this end, it proposes WildClawBench, a benchmark accompanied by an execution-free, task-conditioned evaluation framework. By projecting task scorers into an intent space, the approach quantifies information retention and integrates dedicated ASR models, audio LLMs, context biasing, and LLM-based ontology repair techniques to enhance system robustness. Experimental results demonstrate that the proposed metric improves correlation with downstream task completion rates by 0.246 over WER and CER, offering a more reliable evaluation paradigm for speech-driven agents.
📝 Abstract
Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or targets before the agent begins reasoning, while conventional ASR metrics do not directly measure whether the information required for successful execution has been preserved. We introduce Talk2Agent, a benchmark for evaluating how effectively voice interfaces convey human-spoken instructions to LLM-based computer-use agents. Talk2Agent builds human-spoken versions of tasks from WildClawBench and OSWorld and evaluates a range of voice interfaces, including dedicated ASR models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because repeatedly executing long-horizon computer-use tasks is costly and stochastic, we further propose an execution-free, task-conditioned evaluation framework that projects the original task grader onto prompt-addressable intentions and measures how much task-relevant information is retained after the voice interface. On WildClawBench, Talk2Agent's execution-free native projection provides a practical, execution-grounded measure of voice-interface quality, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.