AS$^2$D: Accelerating On-Demand Audio Understanding on Mobile Devices

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the serial bottleneck in conventional speculative decoding for mobile audio language models, which arises from its dependence on target prefixes during inference. To overcome this limitation, we propose AS$^2$D, a decoupled audio speculative decoding framework. This work is the first to demonstrate that source-conditioned generation can operate independently of evolving target prefixes. By employing an audio-conditioned drafter, AS$^2$D enables concurrent execution of drafting and verification phases, integrated with a batch verification mechanism and deployed on Android devices via MNN. Experiments conducted across four smartphones reveal that the proposed method improves automatic speech recognition throughput by 42–76%, with only 5.7% of processing windows performing slower than the baseline. The achieved performance closely approximates the theoretical optimum, substantially accelerating on-demand audio understanding on mobile platforms.
📝 Abstract
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS$^2$D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS$^2$D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS$^2$D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS$^2$D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
audio understanding
mobile devices
autoregressive generation
inference acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Audio Language Models
Target-Decoupled Drafting
On-Device Inference
Parallel Generation
🔎 Similar Papers
No similar papers found.