Score
Designs and implements pipelines that encode raw speech/audio into vector representations produced by ASR acoustic encoders (including Whisper) and extracts those encoder embeddings. Builds methods to aggregate variable-length encoder outputs—e.g., temporal pooling or attention pooling—and post-process them into fixed- or variable-length embeddings that capture acoustic characteristics for downstream modeling or analysis.
To address Whisper’s inability to support streaming automatic speech recognition (ASR) due to its non-causal encoder-decoder architecture, this paper proposes a unified two-pass decoding framework. The first pass employs a causal CTC decoder for low-latency recognition, while the second pass leverages the original Whisper decoder for semantic rescoring and refinement. Innovatively integrating Whisper with the U2 architecture, we design a hybrid tokenization mechanism: the CTC branch adopts a compact vocabulary for efficiency, whereas the attention branch retains the full vocabulary to preserve semantic completeness. Built upon the WeNet framework, our approach incorporates causal attention masking and joint CTC-attention modeling. Experiments on LibriSpeech and earnings call meeting datasets demonstrate substantial improvements—reducing word error rate by 12.3% and average latency delay by 380 ms—achieving an optimal trade-off between low latency and high accuracy in streaming ASR.
Existing speech foundation models (e.g., Whisper) are optimized for offline, long-audio processing and suffer from high latency and computational overhead on resource-constrained edge devices—primarily due to fixed-length input modeling, costly token encoding, and irregular beam search. Method: We propose HushSpeech, a novel streaming ASR framework featuring “hush words”—learnable silent audio prompts—integrated with learnable audio segment injection, adaptive dynamic beam pruning, and CPU/GPU heterogeneous pipelined scheduling. Contribution/Results: HushSpeech preserves Whisper’s architectural robustness while dramatically improving efficiency: on ARM platforms, it reduces end-to-end latency by 1.6–4.7×, achieving sub-500-ms per-character response; on a MacBook Air, it sustains ~1 s/character throughput at only 7 W total system power—marking the first demonstration of Whisper-level accuracy in energy-efficient, real-time streaming ASR on lightweight edge devices.
To address the challenge of recognizing highly overlapping multi-speaker speech in both online (low-latency) and offline (high-accuracy) automatic speech recognition (ASR) scenarios, this paper proposes an end-to-end unified framework. Methodologically, it introduces the first deep integration of a single-channel continuous speech separation (CSS) frontend with ASR; designs a dual-model collaborative architecture—comprising a Conformer Transducer and a Seq2Seq model—alongside segment-wise serialized output training (segSOT) to jointly enhance robustness to speaker overlap and transcription readability. The approach achieves streaming latency under 300 ms while significantly reducing offline word error rate (WER). Experimental results demonstrate the feasibility and state-of-the-art performance of end-to-end multi-speaker ASR for real-world applications such as live captioning and meeting summarization.
This work proposes a compact and efficient audio–text embedding approach that effectively leverages the Whisper encoder, addressing the limitations of existing methods which underutilize Whisper and suffer from high-dimensional embeddings and low efficiency. The method introduces a learnable global token into the Whisper audio encoder and jointly trains it with a text encoder, employing a two-stage training strategy combined with Matryoshka loss. This is the first framework to successfully integrate Whisper into audio–text embedding learning. The resulting model achieves state-of-the-art performance while enabling an 8× compression of embedding dimensions, delivering strong results on audio–text retrieval, multiple-choice question answering on AIR-Bench, and zero-shot classification tasks. Notably, it substantially reduces storage and computational costs with minimal performance degradation.
This work addresses the limited understanding of internal representation mechanisms in automatic speech recognition (ASR) models by introducing sparse autoencoders (SAEs) to analyze frame-level encoder embeddings of the Whisper model. For the first time, SAEs are employed to construct a high-dimensional sparse latent space that disentangles the semantic structure embedded within Whisper’s representations. The experiments demonstrate that SAEs effectively extract monosemantic features spanning both linguistic and non-linguistic boundaries and enable controllable cross-lingual interventions. These findings not only validate the feasibility of applying SAEs to interpret audio foundation models but also reveal that Whisper encodes rich, structured linguistic information. This study thus establishes a novel paradigm for enhancing the interpretability of ASR systems through sparse representation learning.
研究分析了四个开源音频语言模型对副语言信息的编码和丢失情况,使用多种方法追踪风格信息从音频编码到最终输出的过程,揭示了当前模型在利用副语言信息方面的局限。
This study addresses the inference latency bottleneck in Whisper’s decoder caused by token-by-token generation, proposing an acoustically conditioned speculative decoding method. Exploiting the inherent redundancy in speech signals where “unwritten words are already spoken,” this work designs a two-tier lightweight draft network conditioned on both audio encodings and decoder states to propose eight candidate tokens in parallel within a single forward pass. As the first work to introduce speculative decoding into encoder-decoder speech models, the proposed method achieves a 3.16× inference speedup on the LibriSpeech dataset while maintaining outputs strictly identical to those of the original model. Furthermore, it preserves high efficiency even in large-batch scenarios.
研究通过移除对WER影响最小的六个编码层并利用无标签单语语音数据蒸馏恢复性能,简化Whisper模型结构以提高效率。
This work addresses the challenges faced by conventional neural speech codecs in low-latency scenarios, where heavy frame-level modeling burdens hinder the simultaneous optimization of reconstruction quality and computational efficiency. To overcome this, the authors propose TiCodec, a novel framework featuring a Time-Invariant Representation Extraction (TIRE) module that disentangles speech into time-varying and time-invariant components, substantially reducing frame-level modeling complexity while enabling streaming processing. By incorporating a Dual-TIRE multi-layer architecture that fuses complementary information from different encoder depths, along with factorized representation learning and a 660-ms chunk-based streaming inference strategy, the method significantly enhances both reconstruction fidelity and speaker similarity. Experimental results demonstrate that TiCodec achieves near non-streaming performance under streaming conditions, making it well-suited for low-latency speech generation systems.
This study investigates the joint impact of latent dimensionality and frame rate in continuous audio encoders on downstream task performance. By employing matched training protocols, frozen-model PCA interventions, and automatic speech recognition (ASR) probing techniques, it systematically evaluates representational disparities across varying width and frame rate configurations. The research reveals the interaction mechanisms between these factors regarding downstream utility, demonstrating that representational organization is more critical than mere reconstruction fidelity. Experimental findings indicate that a moderate latent width paired with a high frame rate optimally benefits ASR performance, while high-dimensional models exhibit performance bottlenecks under specific compression ratios. These results challenge the conventional assumption that higher dimensionality is inherently superior, thereby establishing a new paradigm for encoder design.