build speech systems

Designs, builds, and evaluates systems and pipelines that process spoken audio—covering automatic speech recognition (including end-to-end and hybrid ASR), text-to-speech and neural vocoder models, and related model architectures. This includes audio preprocessing and feature extraction, acoustic model training and tuning, end-to-end speech model development, integration of speech-to-text and text-to-speech components into applications, and quantitative evaluation such as ASR evaluation and speech modulation analysis.

buildspeechsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-3.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

On The Landscape of Spoken Language Models: A Comprehensive Survey

Apr 11, 2025
SA
Siddhant Arora
🏛️ Carnegie Mellon University | National Taiwan University | Toyota Technological Institute at Chicago | Hebrew University of Jerusalem | ENS - PSL | EHESS | CNRS

The field of spoken language models (SLMs) suffers from terminological inconsistency, fragmented paradigms, and an unclear developmental trajectory. Method: We conduct the first systematic literature review on SLMs, constructing a structured knowledge graph covering publications from 2020–2025. Our synthesis integrates speech encoders (e.g., Whisper), text-based LMs (e.g., LLaMA), and multimodal alignment techniques, and introduces the first unified SLM taxonomy—distinguishing *pure speech sequence modeling* from *speech-encoder–text-LM fusion*—while clarifying their architectural designs, training strategies, and evaluation protocols. Contribution: This work establishes a consensus-oriented analytical framework for the field, identifies three core challenges—robustness, cross-lingual generalization, and low-resource adaptation—and provides foundational theory and a research roadmap guiding the evolution of SLMs from task-specific models toward general-purpose spoken language processing systems.

Categorizing SLMs by architecture, training, and evaluation methodsIdentifying key challenges and future directions in SLM researchSurveying diverse spoken language models (SLMs) for universal speech processing

Recent Advances in Speech Language Models: A Survey

Oct 01, 2024
WC
Wenqian Cui
🏛️ Chinese University of Hong Kong | Lightspeed Studios | Tencent | National University of Singapore | AI Lab

To address modality distortion, error propagation, and high latency inherent in ASR-TTS cascaded architectures for speech recognition and understanding, this paper presents a systematic survey of recent advances in end-to-end Speech Language Models (SpeechLMs). We propose the first unified analytical framework covering architectural design (speech encoder-decoder), representation schemes (discrete vs. continuous acoustic units), training paradigms (multi-stage pretraining plus instruction tuning), cross-modal alignment mechanisms, and speech-specific evaluation benchmarks. Innovatively, we construct a structured knowledge graph and an open-source resource repository (hosted on GitHub), harmonizing terminology and evaluation standards across the field. Our work delivers both theoretical guidance and practical infrastructure for SpeechLM research, significantly advancing the development of speech-native large language models.

End-to-End ProcessingLanguage UnderstandingSpeech Recognition

Must-Read Papers

Most classic and influential ideas
View more

To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.

Efficiency ImprovementNeural Voice and Audio CodingQuality Evaluation

Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation

Oct 11, 2025
MN
Md. Nayeem
🏛️ National University of Bangladesh | Delhi Technological University (DTU)

This paper surveys the technological evolution of modern automatic speech recognition (ASR), addressing persistent challenges in real-world deployment, including streaming inference, edge-device compatibility, and fairness. Methodologically, it traces architectural advances from hybrid systems to end-to-end models—CTC, RNN-T, Transformer, and Conformer—and critically analyzes the impact of self-supervised pretraining (e.g., wav2vec 2.0, Whisper) and weakly supervised large-scale training in reducing annotation dependency while enhancing data diversity and robustness. Its contributions are threefold: (1) the first unified analysis of practical constraints—latency, computational efficiency, and demographic fairness; (2) a consolidated overview of performance trends across standard benchmarks (LibriSpeech, Switchboard) and standardized evaluation practices (e.g., WER); and (3) a proposed next-generation ASR paradigm centered on “foundation models + weak supervision + diverse data,” offering a theoretical framework and practical guidelines for jointly optimizing efficiency, generalization, and deployability.

Analyzing training paradigm shifts from supervised to self-supervised learningEvaluating deployment challenges including streaming and ethical considerationsSurveying evolution from hybrid to end-to-end neural ASR architectures

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Sep 30, 2024
KL
Ke-Han Lu
🏛️ National Taiwan University | NVIDIA

Speech-language models struggle to jointly optimize speech understanding and textual capability preservation when labeled speech instruction data is scarce. Method: We propose an end-to-end paradigm requiring zero speech instruction data. It aligns a pre-trained speech model with a large language model (LLM) via cross-modal alignment to automatically synthesize high-quality speech-text pairs, thereby injecting paralinguistic understanding; concurrently, the LLM’s parameters are frozen, and only a lightweight speech adapter is introduced to prevent textual capability degradation. Contribution/Results: This work pioneers speech-language co-modeling without any speech instruction fine-tuning data, while supporting joint optimization of complex textual instructions (e.g., chain-of-thought reasoning, format control) and speech understanding. Our approach achieves state-of-the-art performance on Dynamic-SUPERB and AIR-Bench-Chat, significantly reducing reliance on manual speech annotation.

Catastrophic ForgettingSpeech Language ModelingUnsupervised Learning

A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models

Aug 13, 2025
NB
Noureldin Bayoumi
🏛️ RWTH Aachen University | AppTek.ai

Existing ASR model combination methods suffer from evaluation bias due to heterogeneous decoding paradigms (e.g., autoregressive vs. non-autoregressive, frame-level vs. sequence-level outputs) and incompatible label units, hindering fair cross-architecture comparison. Method: This work systematically investigates performance differences and complementarity among four dominant ASR architectures—Attention-based, CTC, Factored Hybrid, and Transducer models—and proposes a unified two-stage decoding framework: (1) independent N-best generation per model; (2) sequence-level log-linear score fusion followed by beam-rescoring for robustness. Contribution/Results: The framework eliminates architectural biases, enabling fair, unit-agnostic evaluation across paradigms. On LibriSpeech 960h, ensemble decoding achieves a 12.3% relative WER reduction over the best single-model baseline, demonstrating the effectiveness and generalizability of heterogeneous ASR architecture collaboration.

Compare model combination across popular ASR architecturesEnsure consistent comparisons in system combination resultsLeverage complementary strengths of different models in search space

ASR systems face a fundamental trade-off between low latency and high accuracy in real-time interpretation scenarios. This paper investigates the impact of audio segmentation strategies on the latency and performance of end-to-end ASR models, proposing a feedback-driven dynamic segmentation algorithm. Unlike fixed-interval segmentation—which minimizes latency but substantially degrades WER—or VAD-based segmentation—which optimizes WER at the cost of maximal latency—our approach dynamically adjusts segment boundaries using real-time recognition feedback to balance both objectives. Experimental results show that our method reduces end-to-end latency by 1.5–2 seconds relative to VAD-based segmentation while increasing WER by only 2–4 percentage points. The proposed feedback-based segmentation is empirically validated on real-time transcription tasks, demonstrating robustness and practicality. This work establishes a deployable paradigm for low-latency, high-fidelity end-to-end ASR systems.

Addressing latency mismatch between ASR transcription and human interpretationMeasuring ASR system delay for real-time interpretation scenariosValidating ASR usability in live settings requiring immediate translation

Latest Papers

What's happening recently
View more

This work proposes GPA, a general-purpose audio model that unifies speech synthesis, automatic speech recognition, and voice conversion within a single autoregressive Transformer architecture—addressing the fragmentation, poor scalability, and limited generalization of traditional task-specific speech systems. By leveraging a shared discrete speech token space and an instruction-driven mechanism, GPA enables zero-architecture-modification task switching. The model employs multi-task joint training and a high-throughput inference pipeline, facilitating lightweight deployment. Experimental results demonstrate that GPA achieves competitive performance across multiple tasks, with its 0.3B-parameter variant particularly well-suited for low-latency, resource-constrained edge scenarios.

autoregressive transformersspeech recognitionspeech synthesis

This work addresses the challenge of effectively leveraging plain text data to enhance the performance of encoder-centric end-to-end automatic speech recognition (ASR) systems. The authors propose a novel approach that integrates modality alignment and dynamic downsampling to enable the encoder to directly produce token-level representations, replacing the conventional large decoder with a compact “large-encoder, small-decoder” architecture. Key innovations include simple yet effective strategies such as stochastic duration modeling. Evaluated on LibriSpeech, the method achieves substantial gains in both recognition accuracy and inference speed, matching or surpassing more complex state-of-the-art systems while significantly streamlining the training pipeline and overall model design. All code and training recipes are publicly released.

encoder-dominated modelsLibriSpeechmodality matching

Discrete audio representations in speech language models often degrade downstream task performance due to information loss. To address this, this work proposes a hybrid discrete-continuous modeling approach that jointly represents speech using temporally compressed discrete tokens and dimensionality-reduced continuous residuals. The method introduces a novel encoder-decoder architecture incorporating fusion-focused modulation and a hybrid Transformer design, enabling autoregressive inference in the discrete domain while simultaneously leveraging non-autoregressive prediction and continuous residual upsampling. This approach achieves the first effective integration of discrete and continuous representations, substantially reducing the number of autoregressive steps while preserving speaker characteristics and fine-grained acoustic details. Experimental results demonstrate clear performance gains over purely discrete baselines.

discrete audio representationsinformation lossLarge Language Models

This work addresses the limitations of existing audio segmentation approaches, which heavily rely on textual transcriptions while neglecting intrinsic audio signals, the impact of ASR errors, and transcription-free evaluation protocols. To overcome these issues, the authors propose AudioSeg, a purely audio-based segmentation model, and conduct a systematic comparison among text-based models, acoustic features, AudioSeg, and multimodal large language models (MLLMs). They further introduce a novel transcription-free evaluation framework based on temporal alignment. Experimental results demonstrate that AudioSeg significantly outperforms text-dependent methods, with silent pauses emerging as the strongest acoustic cue. Although MLLMs are constrained by context length limitations, they show promise on short audio segments. This study establishes the first transcription-independent evaluation benchmark for audio segmentation and provides in-depth analysis of the interplay among transcription quality, acoustic properties, and model performance.

acoustic featuresASR errorsaudio chaptering

This study addresses the limitations of conventional automatic speech recognition (ASR) systems, which are constrained by short-term context modeling and the i.i.d. assumption, hindering effective exploitation of long-range acoustic and linguistic information. For the first time, this work empirically demonstrates that incorporating contextual information spanning up to 21.8 minutes significantly improves recognition performance in end-to-end attention-based ASR models, achieving a relative word error rate reduction of up to 14.2%. Through systematic evaluation of key architectural factors—including positional encoding schemes, model depth, and width—the study elucidates their impact on modeling ultra-long sequences (up to one hour), thereby providing both an effective architecture and empirical foundation for long-context speech recognition.

automatic speech recognitioncontext utilizationlong-context modeling

Hot Scholars

HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
ZW

Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing