extract asr embeddings

Designs and implements pipelines that encode raw speech/audio into vector representations produced by ASR acoustic encoders (including Whisper) and extracts those encoder embeddings. Builds methods to aggregate variable-length encoder outputs—e.g., temporal pooling or attention pooling—and post-process them into fixed- or variable-length embeddings that capture acoustic characteristics for downstream modeling or analysis.

extractasrembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Jun 13, 2025
HZ
Haoran Zhou
🏛️ Bloomberg | WeNet Open Source Community

To address Whisper’s inability to support streaming automatic speech recognition (ASR) due to its non-causal encoder-decoder architecture, this paper proposes a unified two-pass decoding framework. The first pass employs a causal CTC decoder for low-latency recognition, while the second pass leverages the original Whisper decoder for semantic rescoring and refinement. Innovatively integrating Whisper with the U2 architecture, we design a hybrid tokenization mechanism: the CTC branch adopts a compact vocabulary for efficiency, whereas the attention branch retains the full vocabulary to preserve semantic completeness. Built upon the WeNet framework, our approach incorporates causal attention masking and joint CTC-attention modeling. Experiments on LibriSpeech and earnings call meeting datasets demonstrate substantial improvements—reducing word error rate by 12.3% and average latency delay by 380 ms—achieving an optimal trade-off between low latency and high accuracy in streaming ASR.

Adapt Whisper for streaming speech recognitionImplement two-pass decoding for partial transcriptsImprove data efficiency with hybrid tokenizer

Efficient Whisper on Streaming Speech

Dec 15, 2024
RW
Rongxiang Wang
🏛️ University of Virginia | Independent researcher

Existing speech foundation models (e.g., Whisper) are optimized for offline, long-audio processing and suffer from high latency and computational overhead on resource-constrained edge devices—primarily due to fixed-length input modeling, costly token encoding, and irregular beam search. Method: We propose HushSpeech, a novel streaming ASR framework featuring “hush words”—learnable silent audio prompts—integrated with learnable audio segment injection, adaptive dynamic beam pruning, and CPU/GPU heterogeneous pipelined scheduling. Contribution/Results: HushSpeech preserves Whisper’s architectural robustness while dramatically improving efficiency: on ARM platforms, it reduces end-to-end latency by 1.6–4.7×, achieving sub-500-ms per-character response; on a MacBook Air, it sustains ~1 s/character throughput at only 7 W total system power—marking the first demonstration of Whisper-level accuracy in energy-efficient, real-time streaming ASR on lightweight edge devices.

High resource usage on client devicesInefficient long input encoding and decodingReal-time streaming speech processing limitations

To address the challenge of recognizing highly overlapping multi-speaker speech in both online (low-latency) and offline (high-accuracy) automatic speech recognition (ASR) scenarios, this paper proposes an end-to-end unified framework. Methodologically, it introduces the first deep integration of a single-channel continuous speech separation (CSS) frontend with ASR; designs a dual-model collaborative architecture—comprising a Conformer Transducer and a Seq2Seq model—alongside segment-wise serialized output training (segSOT) to jointly enhance robustness to speaker overlap and transcription readability. The approach achieves streaming latency under 300 ms while significantly reducing offline word error rate (WER). Experimental results demonstrate the feasibility and state-of-the-art performance of end-to-end multi-speaker ASR for real-world applications such as live captioning and meeting summarization.

Balancing latency and accuracy in streaming and offline ASREnhancing multi-talker transcription readability with segment-based SOTImproving accuracy in overlapping speech with CSS and E2E systems

This work proposes a compact and efficient audio–text embedding approach that effectively leverages the Whisper encoder, addressing the limitations of existing methods which underutilize Whisper and suffer from high-dimensional embeddings and low efficiency. The method introduces a learnable global token into the Whisper audio encoder and jointly trains it with a text encoder, employing a two-stage training strategy combined with Matryoshka loss. This is the first framework to successfully integrate Whisper into audio–text embedding learning. The resulting model achieves state-of-the-art performance while enabling an 8× compression of embedding dimensions, delivering strong results on audio–text retrieval, multiple-choice question answering on AIR-Bench, and zero-shot classification tasks. Notably, it substantially reduces storage and computational costs with minimal performance degradation.

audio-text embeddingcompact representationretrieval performance

This work addresses the limited understanding of internal representation mechanisms in automatic speech recognition (ASR) models by introducing sparse autoencoders (SAEs) to analyze frame-level encoder embeddings of the Whisper model. For the first time, SAEs are employed to construct a high-dimensional sparse latent space that disentangles the semantic structure embedded within Whisper’s representations. The experiments demonstrate that SAEs effectively extract monosemantic features spanning both linguistic and non-linguistic boundaries and enable controllable cross-lingual interventions. These findings not only validate the feasibility of applying SAEs to interpret audio foundation models but also reveal that Whisper encodes rich, structured linguistic information. This study thus establishes a novel paradigm for enhancing the interpretability of ASR systems through sparse representation learning.

Automatic Speech RecognitionMechanistic InterpretabilitySparse Autoencoders

Latest Papers

What's happening recently
View more

This study addresses the inference latency bottleneck in Whisper’s decoder caused by token-by-token generation, proposing an acoustically conditioned speculative decoding method. Exploiting the inherent redundancy in speech signals where “unwritten words are already spoken,” this work designs a two-tier lightweight draft network conditioned on both audio encodings and decoder states to propose eight candidate tokens in parallel within a single forward pass. As the first work to introduce speculative decoding into encoder-decoder speech models, the proposed method achieves a 3.16× inference speedup on the LibriSpeech dataset while maintaining outputs strictly identical to those of the original model. Furthermore, it preserves high efficiency even in large-batch scenarios.

autoregressive generationdecoding latencyinference speed

This work addresses the challenges faced by conventional neural speech codecs in low-latency scenarios, where heavy frame-level modeling burdens hinder the simultaneous optimization of reconstruction quality and computational efficiency. To overcome this, the authors propose TiCodec, a novel framework featuring a Time-Invariant Representation Extraction (TIRE) module that disentangles speech into time-varying and time-invariant components, substantially reducing frame-level modeling complexity while enabling streaming processing. By incorporating a Dual-TIRE multi-layer architecture that fuses complementary information from different encoder depths, along with factorized representation learning and a 660-ms chunk-based streaming inference strategy, the method significantly enhances both reconstruction fidelity and speaker similarity. Experimental results demonstrate that TiCodec achieves near non-streaming performance under streaming conditions, making it well-suited for low-latency speech generation systems.

low-latency speech processingneural speech codecsspeech generation

This study investigates the joint impact of latent dimensionality and frame rate in continuous audio encoders on downstream task performance. By employing matched training protocols, frozen-model PCA interventions, and automatic speech recognition (ASR) probing techniques, it systematically evaluates representational disparities across varying width and frame rate configurations. The research reveals the interaction mechanisms between these factors regarding downstream utility, demonstrating that representational organization is more critical than mere reconstruction fidelity. Experimental findings indicate that a moderate latent width paired with a high frame rate optimally benefits ASR performance, while high-dimensional models exhibit performance bottlenecks under specific compression ratios. These results challenge the conventional assumption that higher dimensionality is inherently superior, thereby establishing a new paradigm for encoder design.

audio compressioncontinuous audio encodersdownstream performance

Hot Scholars

JC

Junjie Cao

School of Mathematical Sciences, Dalian University of Technology
Computer GraphicsComputer VisionMachine Learning
PM

Petr Motlicek

Idiap Research Institute
Artificial intelligencespeech and signal processingmachine learning
PP

Pablo Peso Parada

AI Researcher - Samsung Research UK
signal processingmachine learningopen source hardwareaudio