🤖 AI Summary
This study addresses the response latency caused by serial tool invocation in voice assistants by proposing a speculative execution architecture based on partial ASR hypotheses. The method predicts and pre-executes tool calls during speech recognition, injects the cached results into the LLM context, and introduces a rule-based verification engine to handle user self-corrections and rollbacks, ensuring worst-case performance no worse than the baseline. Empirical evaluations demonstrate that the system reduces the median time-to-first-audio from 5.79s to 4.60s while narrowing the standard deviation to 2.81s, significantly improving both end-to-end response speed and predictability.
📝 Abstract
Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.