Score
Designs, builds, or analyzes end-to-end CTC-based speech recognition systems that map raw audio waveforms to phonetic character sequences (e.g., IPA), producing whitespace-insensitive phonetic transcripts. Work includes specifying architectures and CTC training/decoding pipelines, choosing phonetic output representations, and optimizing models for constrained parameter budgets.
This work addresses the challenge of achieving both low-latency streaming synthesis and high audio quality in large language model (LLM)-based text-to-speech (TTS) systems, which typically rely on cumbersome GMM-HMM forced alignment tools like MFA. The authors propose CTC-TTS, the first framework to integrate a CTC-based neural aligner into LLM-based TTS, replacing the conventional MFA pipeline. A dual-token interleaving strategy is introduced to more accurately model text–speech alignment. Two variants are developed: CTC-TTS-L enhances audio quality through token concatenation, while CTC-TTS-F reduces latency via embedding stacking. Experimental results demonstrate that the proposed approach outperforms both MFA-based and fixed-interleaving baselines in streaming and zero-shot settings, significantly improving speech naturalness and response speed.
This study addresses the challenge of deploying large-scale multilingual IPA transcription models on resource-constrained devices by proposing a lightweight speech recognition architecture. The method constructs a compact backbone network using an E-Branchformer encoder integrated with rotary position encoding, and introduces self-conditioned CTC alongside data augmentation consistency regularization for synergistic training optimization. Experimental results demonstrate that the proposed model substantially reduces parameter count while significantly improving recognition accuracy, achieving an IPA character error rate of 4.47%—a 22.3% relative reduction over the baseline. Furthermore, it outperforms Conformer models of comparable scale, effectively facilitating efficient on-device deployment.
This work addresses the challenge of effectively leveraging plain text data to enhance the performance of encoder-centric end-to-end automatic speech recognition (ASR) systems. The authors propose a novel approach that integrates modality alignment and dynamic downsampling to enable the encoder to directly produce token-level representations, replacing the conventional large decoder with a compact “large-encoder, small-decoder” architecture. Key innovations include simple yet effective strategies such as stochastic duration modeling. Evaluated on LibriSpeech, the method achieves substantial gains in both recognition accuracy and inference speed, matching or surpassing more complex state-of-the-art systems while significantly streamlining the training pipeline and overall model design. All code and training recipes are publicly released.
To address the challenges of long-range contextual dependencies and language-specific constraints in universal phoneme recognition, this paper proposes a context-agnostic and language-agnostic universal phoneme encoding method. The approach extracts acoustic features from fixed-width, 120-ms short windows independently—eliminating reliance on sequential modeling. A lightweight single-window model architecture is designed to explicitly capture cross-lingual commonalities in acoustic patterns. Supervised and self-supervised learning objectives are jointly optimized across multilingual data, enabling zero-shot cross-lingual transfer. Evaluated on multilingual phoneme recognition benchmarks, the method achieves competitive performance. Notably, it significantly outperforms existing context-free approaches on the UCLA zero-shot cross-lingual evaluation, demonstrating superior generalization. This work is the first to empirically validate the feasibility and strong transferability of short-window independent encoding for universal phoneme representation learning.
To address the challenges of poor speech-text modality alignment and weak cross-domain generalization in decoder-only end-to-end ASR, this paper proposes a CTC-compressor-driven joint training framework. Methodologically, it introduces (1) a novel bidirectional modality matching mechanism that integrates CTC compression, dynamic forced spike alignment, and CTC-inspired embeddings—enabling efficient audio-text fusion without explicit duration modeling; and (2) a lightweight modality adapter to enhance cross-domain robustness. Evaluated on LibriSpeech and TED-LIUM2, the approach achieves state-of-the-art performance among decoder-only models of comparable size, with significant improvements in noise robustness, long-form speech recognition, and cross-domain generalization. Furthermore, the study systematically characterizes optimal configurations of the CTC compressor under boundary and noisy conditions, providing principled guidance for its deployment in challenging acoustic scenarios.
This work proposes BranchShine, a lightweight end-to-end model for multilingual speech-to-IPA transcription that operates directly on raw audio and comprises only 33 million parameters. The architecture features a compact convolutional frontend followed by a 19-layer E-Branchformer encoder enhanced with rotary position embeddings (RoPE), trained with CTC loss. Evaluated on a diverse test set of 16,660 utterances spanning 41 languages, BranchShine achieves a space-insensitive IPA character error rate of 9.19%, outperforming the 575-million-parameter PhoneticXEUS baseline (9.78%). This demonstrates that substantial model compression can be achieved without sacrificing performance, offering an efficient and compact solution for character-level IPA transcription.
This work proposes a phoneme-level end-to-end automatic speech recognition (ASR) approach for Vietnamese that explicitly incorporates syllabic phonological constraints during decoding to generate valid syllables from a compact phoneme inventory. Unlike conventional ASR systems that rely on character- or subword-level units and require large vocabularies, the proposed method leverages the intricate syllable structure of Vietnamese without needing additional training data or pretrained models. Evaluated on the LSVSC and UIT-ViMD benchmarks, the system outperforms strong baselines such as PhoWhisper and Wav2Vec2, achieving higher recognition accuracy with a substantially reduced vocabulary size. Notably, it demonstrates robust performance across multiple dialects, highlighting its effectiveness in handling linguistic variation inherent in Vietnamese speech.
This study addresses the challenge of acquiring novel vocabulary during inference in automatic speech recognition (ASR) systems. To this end, it proposes a test-time adaptation framework that leverages unannotated data to learn contextual representations and spellings of new words while keeping both the acoustic and language models frozen. The core innovation lies in equivalently reformulating the CTC-weighted log-likelihood ratio as a Kullback–Leibler divergence minimization objective, accompanied by theoretical guarantees establishing an upper bound on the total variation distance. Experimental evaluations on LibriSpeech and a dysarthric speech dataset demonstrate that the proposed approach significantly reduces character error rates for repeated out-of-vocabulary words by 14.97% and 6.67%, respectively, thereby effectively enhancing the system’s capacity for vocabulary expansion.
This work addresses the challenge of diacritic omission in Arabic speech transcription, which obscures fine-grained phonetic distinctions and hinders accurate modeling. The authors propose a CTC-based non-autoregressive approach for diacritic restoration that constructs a character-level diacritic lattice and enforces hard constraints during decoding to retain only linguistically valid diacritized sequences. By explicitly leveraging acoustic information and drastically reducing the decoding search space through these constraints, the method achieves superior accuracy while maintaining computational efficiency. Evaluated on the ArVoice and ClArTTS benchmarks, the proposed approach significantly lowers diacritic error rates compared to existing multimodal baselines, demonstrating both effectiveness and practicality for real-world deployment.
为解决发音评估中声学模型难以同时提供识别和分割证据的问题,提出结合有序子音素状态与最优时间传输分类的拓扑感知帧级声学模型。