Score
Designs, builds, and evaluates systems and pipelines that process spoken audio—covering automatic speech recognition (including end-to-end and hybrid ASR), text-to-speech and neural vocoder models, and related model architectures. This includes audio preprocessing and feature extraction, acoustic model training and tuning, end-to-end speech model development, integration of speech-to-text and text-to-speech components into applications, and quantitative evaluation such as ASR evaluation and speech modulation analysis.
Transformers face inherent limitations in speech processing—including modeling long-range temporal dependencies, computational redundancy, and low data efficiency—across diverse tasks. Method: This work establishes a unified analytical framework that systematically integrates self-attention mechanisms, positional encodings, and the pretraining-finetuning paradigm, while incorporating MFCC/wav2vec feature representations, multi-scale temporal modeling, and cross-modal alignment techniques. Contribution/Results: The framework comprehensively covers seven major speech tasks—ASR, text-to-speech, speech translation, paralinguistic analysis, enhancement, dialogue systems, and multimodal processing—identifying 12 recurring technical challenges and surveying over 30 representative models. It yields a structured, taxonomy-driven survey that pinpoints core bottlenecks (e.g., temporal modeling bias) and proposes reproducible improvement pathways, serving as an authoritative reference and technical roadmap for Transformer-based speech research.
The field of spoken language models (SLMs) suffers from terminological inconsistency, fragmented paradigms, and an unclear developmental trajectory. Method: We conduct the first systematic literature review on SLMs, constructing a structured knowledge graph covering publications from 2020–2025. Our synthesis integrates speech encoders (e.g., Whisper), text-based LMs (e.g., LLaMA), and multimodal alignment techniques, and introduces the first unified SLM taxonomy—distinguishing *pure speech sequence modeling* from *speech-encoder–text-LM fusion*—while clarifying their architectural designs, training strategies, and evaluation protocols. Contribution: This work establishes a consensus-oriented analytical framework for the field, identifies three core challenges—robustness, cross-lingual generalization, and low-resource adaptation—and provides foundational theory and a research roadmap guiding the evolution of SLMs from task-specific models toward general-purpose spoken language processing systems.
To address modality distortion, error propagation, and high latency inherent in ASR-TTS cascaded architectures for speech recognition and understanding, this paper presents a systematic survey of recent advances in end-to-end Speech Language Models (SpeechLMs). We propose the first unified analytical framework covering architectural design (speech encoder-decoder), representation schemes (discrete vs. continuous acoustic units), training paradigms (multi-stage pretraining plus instruction tuning), cross-modal alignment mechanisms, and speech-specific evaluation benchmarks. Innovatively, we construct a structured knowledge graph and an open-source resource repository (hosted on GitHub), harmonizing terminology and evaluation standards across the field. Our work delivers both theoretical guidance and practical infrastructure for SpeechLM research, significantly advancing the development of speech-native large language models.
To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.
This paper surveys the technological evolution of modern automatic speech recognition (ASR), addressing persistent challenges in real-world deployment, including streaming inference, edge-device compatibility, and fairness. Methodologically, it traces architectural advances from hybrid systems to end-to-end models—CTC, RNN-T, Transformer, and Conformer—and critically analyzes the impact of self-supervised pretraining (e.g., wav2vec 2.0, Whisper) and weakly supervised large-scale training in reducing annotation dependency while enhancing data diversity and robustness. Its contributions are threefold: (1) the first unified analysis of practical constraints—latency, computational efficiency, and demographic fairness; (2) a consolidated overview of performance trends across standard benchmarks (LibriSpeech, Switchboard) and standardized evaluation practices (e.g., WER); and (3) a proposed next-generation ASR paradigm centered on “foundation models + weak supervision + diverse data,” offering a theoretical framework and practical guidelines for jointly optimizing efficiency, generalization, and deployability.
Speech-language models struggle to jointly optimize speech understanding and textual capability preservation when labeled speech instruction data is scarce. Method: We propose an end-to-end paradigm requiring zero speech instruction data. It aligns a pre-trained speech model with a large language model (LLM) via cross-modal alignment to automatically synthesize high-quality speech-text pairs, thereby injecting paralinguistic understanding; concurrently, the LLM’s parameters are frozen, and only a lightweight speech adapter is introduced to prevent textual capability degradation. Contribution/Results: This work pioneers speech-language co-modeling without any speech instruction fine-tuning data, while supporting joint optimization of complex textual instructions (e.g., chain-of-thought reasoning, format control) and speech understanding. Our approach achieves state-of-the-art performance on Dynamic-SUPERB and AIR-Bench-Chat, significantly reducing reliance on manual speech annotation.
Existing ASR model combination methods suffer from evaluation bias due to heterogeneous decoding paradigms (e.g., autoregressive vs. non-autoregressive, frame-level vs. sequence-level outputs) and incompatible label units, hindering fair cross-architecture comparison. Method: This work systematically investigates performance differences and complementarity among four dominant ASR architectures—Attention-based, CTC, Factored Hybrid, and Transducer models—and proposes a unified two-stage decoding framework: (1) independent N-best generation per model; (2) sequence-level log-linear score fusion followed by beam-rescoring for robustness. Contribution/Results: The framework eliminates architectural biases, enabling fair, unit-agnostic evaluation across paradigms. On LibriSpeech 960h, ensemble decoding achieves a 12.3% relative WER reduction over the best single-model baseline, demonstrating the effectiveness and generalizability of heterogeneous ASR architecture collaboration.
ASR systems face a fundamental trade-off between low latency and high accuracy in real-time interpretation scenarios. This paper investigates the impact of audio segmentation strategies on the latency and performance of end-to-end ASR models, proposing a feedback-driven dynamic segmentation algorithm. Unlike fixed-interval segmentation—which minimizes latency but substantially degrades WER—or VAD-based segmentation—which optimizes WER at the cost of maximal latency—our approach dynamically adjusts segment boundaries using real-time recognition feedback to balance both objectives. Experimental results show that our method reduces end-to-end latency by 1.5–2 seconds relative to VAD-based segmentation while increasing WER by only 2–4 percentage points. The proposed feedback-based segmentation is empirically validated on real-time transcription tasks, demonstrating robustness and practicality. This work establishes a deployable paradigm for low-latency, high-fidelity end-to-end ASR systems.
This work proposes GPA, a general-purpose audio model that unifies speech synthesis, automatic speech recognition, and voice conversion within a single autoregressive Transformer architecture—addressing the fragmentation, poor scalability, and limited generalization of traditional task-specific speech systems. By leveraging a shared discrete speech token space and an instruction-driven mechanism, GPA enables zero-architecture-modification task switching. The model employs multi-task joint training and a high-throughput inference pipeline, facilitating lightweight deployment. Experimental results demonstrate that GPA achieves competitive performance across multiple tasks, with its 0.3B-parameter variant particularly well-suited for low-latency, resource-constrained edge scenarios.
This work addresses the challenge of effectively leveraging plain text data to enhance the performance of encoder-centric end-to-end automatic speech recognition (ASR) systems. The authors propose a novel approach that integrates modality alignment and dynamic downsampling to enable the encoder to directly produce token-level representations, replacing the conventional large decoder with a compact “large-encoder, small-decoder” architecture. Key innovations include simple yet effective strategies such as stochastic duration modeling. Evaluated on LibriSpeech, the method achieves substantial gains in both recognition accuracy and inference speed, matching or surpassing more complex state-of-the-art systems while significantly streamlining the training pipeline and overall model design. All code and training recipes are publicly released.
Discrete audio representations in speech language models often degrade downstream task performance due to information loss. To address this, this work proposes a hybrid discrete-continuous modeling approach that jointly represents speech using temporally compressed discrete tokens and dimensionality-reduced continuous residuals. The method introduces a novel encoder-decoder architecture incorporating fusion-focused modulation and a hybrid Transformer design, enabling autoregressive inference in the discrete domain while simultaneously leveraging non-autoregressive prediction and continuous residual upsampling. This approach achieves the first effective integration of discrete and continuous representations, substantially reducing the number of autoregressive steps while preserving speaker characteristics and fine-grained acoustic details. Experimental results demonstrate clear performance gains over purely discrete baselines.
This work addresses the limitations of existing audio segmentation approaches, which heavily rely on textual transcriptions while neglecting intrinsic audio signals, the impact of ASR errors, and transcription-free evaluation protocols. To overcome these issues, the authors propose AudioSeg, a purely audio-based segmentation model, and conduct a systematic comparison among text-based models, acoustic features, AudioSeg, and multimodal large language models (MLLMs). They further introduce a novel transcription-free evaluation framework based on temporal alignment. Experimental results demonstrate that AudioSeg significantly outperforms text-dependent methods, with silent pauses emerging as the strongest acoustic cue. Although MLLMs are constrained by context length limitations, they show promise on short audio segments. This study establishes the first transcription-independent evaluation benchmark for audio segmentation and provides in-depth analysis of the interplay among transcription quality, acoustic properties, and model performance.
This study addresses the limitations of conventional automatic speech recognition (ASR) systems, which are constrained by short-term context modeling and the i.i.d. assumption, hindering effective exploitation of long-range acoustic and linguistic information. For the first time, this work empirically demonstrates that incorporating contextual information spanning up to 21.8 minutes significantly improves recognition performance in end-to-end attention-based ASR models, achieving a relative word error rate reduction of up to 14.2%. Through systematic evaluation of key architectural factors—including positional encoding schemes, model depth, and width—the study elucidates their impact on modeling ultra-long sequences (up to one hour), thereby providing both an effective architecture and empirical foundation for long-context speech recognition.