Score
Designs, builds, and evaluates systems that transcribe spoken-language audio into text, including front-end signal processing, acoustic and language modeling (traditional and end-to-end), decoding/beam search and runtime/streaming architectures. Works on dataset creation and annotation, training and adaptation for speakers and noise conditions, inference optimization, and evaluation using transcription accuracy and error-rate metrics.
The field of spoken language models (SLMs) suffers from terminological inconsistency, fragmented paradigms, and an unclear developmental trajectory. Method: We conduct the first systematic literature review on SLMs, constructing a structured knowledge graph covering publications from 2020–2025. Our synthesis integrates speech encoders (e.g., Whisper), text-based LMs (e.g., LLaMA), and multimodal alignment techniques, and introduces the first unified SLM taxonomy—distinguishing *pure speech sequence modeling* from *speech-encoder–text-LM fusion*—while clarifying their architectural designs, training strategies, and evaluation protocols. Contribution: This work establishes a consensus-oriented analytical framework for the field, identifies three core challenges—robustness, cross-lingual generalization, and low-resource adaptation—and provides foundational theory and a research roadmap guiding the evolution of SLMs from task-specific models toward general-purpose spoken language processing systems.
This study investigates the trade-off between efficiency and accuracy in semi-automatic transcription for spoken language corpus construction. Through a two-stage experiment, it compares the performance of expert and novice transcribers on three types of Italian conversational data under both manual and ASR-assisted conditions. The work proposes an integrated analytical framework combining word-level alignment, quality evaluation metrics, and statistical modeling to systematically quantify behavioral differences across transcription workflows. Results demonstrate that ASR substantially increases transcription speed, yet its impact on accuracy varies depending on dialogue type, transcriber expertise, and workflow configuration. The findings provide empirical support for the development of the KIParla corpus, showing that a fine-tuned and optimized semi-automatic pipeline can effectively accelerate annotation while maintaining high transcription quality.
This study addresses the underexplored, high-difficulty ASR task of poetry speech recognition. We introduce the first multi-system benchmark for poetry reading—built upon nearly 10 hours of authentic recordings from PennSound, encompassing diverse acoustic conditions and stylistically rich, prosodic, and colloquial speech. We systematically evaluate eight mainstream commercial and open-source ASR systems, including AWS, Azure, and Whisper. Our contributions are threefold: (1) pioneering the use of poetry corpora for cross-system ASR evaluation; (2) revealing, for the first time, Whisper’s hallucination propensity as strongly correlated with decoding parameters; and (3) proposing a novel multi-dimensional evaluation framework integrating Word Error Rate (WER) and Diarization Error Rate (DER). Results show Rev.ai achieves the best overall performance; Whisper is the top open-source model—provided hallucinations are mitigated via targeted tuning; and AWS excels in speaker diarization. Notably, performance gaps among systems are narrow, underscoring the critical importance of task-specific adaptation.
This work addresses the limited robustness and integrative reasoning capabilities of current speech models in long-form scenarios—such as meeting transcription and spoken document understanding—despite their strong performance on short utterances. To bridge this gap, we introduce LongSpeech, the first large-scale, extensible multitask benchmark for long speech, comprising over 100,000 audio segments averaging ten minutes each. LongSpeech supports diverse tasks including automatic speech recognition, speech translation, summarization, language identification, speaker counting, content disentanglement, and question answering. The benchmark is built from heterogeneous data sources and features multidimensional manual and automatic annotations, standardized evaluation protocols, and a reproducible construction pipeline, offering a unified platform for long-speech research. Preliminary evaluations reveal substantial performance gaps in state-of-the-art models, particularly in cross-task generalization and higher-order reasoning.
Speech-language models struggle to jointly optimize speech understanding and textual capability preservation when labeled speech instruction data is scarce. Method: We propose an end-to-end paradigm requiring zero speech instruction data. It aligns a pre-trained speech model with a large language model (LLM) via cross-modal alignment to automatically synthesize high-quality speech-text pairs, thereby injecting paralinguistic understanding; concurrently, the LLM’s parameters are frozen, and only a lightweight speech adapter is introduced to prevent textual capability degradation. Contribution/Results: This work pioneers speech-language co-modeling without any speech instruction fine-tuning data, while supporting joint optimization of complex textual instructions (e.g., chain-of-thought reasoning, format control) and speech understanding. Our approach achieves state-of-the-art performance on Dynamic-SUPERB and AIR-Bench-Chat, significantly reducing reliance on manual speech annotation.
This work addresses the limitations of existing speech-to-text translation systems, which typically rely on cascaded architectures prone to error propagation, and current SpeechLLMs that fail to support truly real-time streaming translation, resulting in high latency. The paper proposes the first genuinely streaming SpeechLLM architecture, wherein a large language model learns an end-to-end mapping from speech to translated text and dynamically determines output timing—autonomously deciding when sufficient acoustic context is available to generate the next token, without relying on fixed intervals or waiting for complete utterances. By incorporating paralinguistic cues and training on automatically aligned speech–text data, the method achieves translation quality approaching that of non-streaming baselines across multiple languages while maintaining latency within 1–2 seconds, substantially enhancing practicality and responsiveness.
This work addresses the challenge of effectively leveraging plain text data to enhance the performance of encoder-centric end-to-end automatic speech recognition (ASR) systems. The authors propose a novel approach that integrates modality alignment and dynamic downsampling to enable the encoder to directly produce token-level representations, replacing the conventional large decoder with a compact “large-encoder, small-decoder” architecture. Key innovations include simple yet effective strategies such as stochastic duration modeling. Evaluated on LibriSpeech, the method achieves substantial gains in both recognition accuracy and inference speed, matching or surpassing more complex state-of-the-art systems while significantly streamlining the training pipeline and overall model design. All code and training recipes are publicly released.
This work addresses the performance degradation of task-oriented spoken language understanding systems in noisy environments due to automatic speech recognition (ASR) errors, and the high cost of acquiring speech–semantics aligned annotations for adapting speech language models to new tasks. The authors propose CORTIS, a novel framework that, for the first time, enables end-to-end fine-tuning of speech language models using only textual task supervision signals—eliminating the need for task-specific speech annotations. Experiments based on Qwen2.5-Omni demonstrate that CORTIS matches the performance of conventional ASR–LLM cascade systems on three task-oriented spoken datasets under clean conditions, while significantly outperforming them in noisy settings, particularly exhibiting superior robustness in preserving high-level semantic content.
This study addresses the persistent challenges in low-resource speech synthesis—namely poor quality and weak generalization—stemming from scarce authentic corpora and orthographic diversity. To this end, we introduce OpenBibleTTS, the first large-scale multilingual text-to-speech (TTS) benchmark built upon authentic Bible texts and out-of-domain data, spanning 37 low-resource languages. We systematically evaluate state-of-the-art architectures, including Gemini-TTS and EveryVoice, in terms of intelligibility, naturalness, and cross-domain robustness. Our analysis reveals a trade-off between multilingual and monolingual approaches: Gemini-TTS achieves the highest subjective ratings across most languages, whereas monolingual EveryVoice demonstrates superior intelligibility for African languages. All data and models are publicly released to foster fair, reproducible research in low-resource TTS.