Whisper-Flash: Acoustically Conditioned Parallel Drafting for Faster Whisper Decoding

๐Ÿ“… 2026-09-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the inference latency bottleneck in Whisperโ€™s decoder caused by token-by-token generation, proposing an acoustically conditioned speculative decoding method. Exploiting the inherent redundancy in speech signals where โ€œunwritten words are already spoken,โ€ this work designs a two-tier lightweight draft network conditioned on both audio encodings and decoder states to propose eight candidate tokens in parallel within a single forward pass. As the first work to introduce speculative decoding into encoder-decoder speech models, the proposed method achieves a 3.16ร— inference speedup on the LibriSpeech dataset while maintaining outputs strictly identical to those of the original model. Furthermore, it preserves high efficiency even in large-batch scenarios.
๐Ÿ“ Abstract
Whisper is a widely used encoder-decoder model for speech recognition. Its encoder reads an utterance in one parallel pass, but its decoder writes the transcript one token at a time, which dominates inference time. Speculative decoding shortens such loops without changing their output: a small drafter guesses several upcoming tokens, and the original model verifies them all in one forward pass. We present Whisper-Flash, a two-layer drafter built on a property of speech recognition: the words still to be written have already been spoken. It reads Whisper's encoded audio and accepted decoder states and proposes eight tokens in a single forward pass. On the complete LibriSpeech test sets, Whisper-Flash processes $3.16\times/2.85\times$ as much audio per second as greedy decoding with identical outputs, and it remains faster at batch sizes up to 96 and under temperature sampling. Ablations show that direct access to the audio matters most.
Problem

Research questions and friction points this paper is trying to address.

Whisper
speech recognition
decoding latency
autoregressive generation
inference speed
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Draft Model
Speech Recognition
Parallel Drafting
Inference Acceleration
๐Ÿ”Ž Similar Papers
No similar papers found.