train with ctc

Designs and implements sequence models trained with Connectionist Temporal Classification (CTC) loss to map variable-length input sequences to output token sequences without frame-level alignment; this includes computing the CTC objective, optimizing model parameters, handling blank symbols and repeated-label collapsing, batching and numerical stability, and building decoding/beam-search and sequence-level evaluation components for alignment-free training.

trainwithctc

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

CTC, while computationally efficient for automatic speech recognition (ASR), suffers from suboptimal performance due to its sharp output distribution and lack of contextual modeling. To address this, we propose Consistency-Regularized CTC (CR-CTC), the first method to introduce self-consistency regularization into the CTC framework. CR-CTC generates multiple augmented views of mel-spectrogram inputs, applies independent CTC modeling to each view, and enforces consistency among their output distributions via KL-divergence constraints. This yields implicit self-distillation and context-aware temporal masking representation learning, effectively mitigating CTC’s peaky output problem. Crucially, CR-CTC requires no architectural modifications or auxiliary decoders—only a principled loss-function redesign. Evaluated on LibriSpeech, AISHELL-1, and GigaSpeech, CR-CTC consistently outperforms standard CTC, matches or exceeds the accuracy of RNN-T and CTC/attention hybrid systems, and achieves state-of-the-art performance across benchmarks.

Enhances contextual representation learningImproves CTC speech recognitionReduces overfitting in CTC distributions

This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.

automatic speech recognitionencoder frame gridspeech LLMs

Label-Context-Dependent Internal Language Model Estimation for CTC

Jun 06, 2025
ZY
Zijian Yang
🏛️ RWTH Aachen University | AppTek GmbH

Connectionist Temporal Classification (CTC) theoretically assumes conditional independence among output labels, yet empirical evidence suggests that strong encoders implicitly learn context-dependent internal language models (ILMs). Method: This work formally models and empirically validates this implicit contextual dependency for the first time; proposes a label-level knowledge distillation framework to unsupervisedly extract context-aware ILMs from CTC decoders; and introduces two novel strategies—smooth regularization and label-level contextual modeling—to relax the conventional independent-label assumption. Contribution/Results: On the cross-domain TED-LIUM benchmark, the proposed approach reduces word error rate (WER) by over 13% compared to shallow fusion, significantly outperforming context-agnostic priors. These results robustly demonstrate CTC’s capacity to learn strong implicit ILMs, challenging the standard independence assumption and enabling more effective integration of contextual knowledge in end-to-end speech recognition.

Estimating context-dependent ILM in CTC modelsEvaluating ILM impact on cross-domain speech recognitionImproving CTC performance via knowledge distillation

Temporal Chunking Enhances Recognition of Implicit Sequential Patterns

May 31, 2025
JD
Jayanta Dey
🏛️ University of Texas at San Antonio | University of Rochester

Traditional RNNs struggle to capture implicit temporal patterns spanning multiple time scales. Method: Inspired by neuroscience, we propose a “wake–sleep” two-phase learning framework. During the offline “sleep phase,” a graph-community-detection–based temporal chunking mechanism automatically generates context-aware, structured memory units—enabling long-sequence compression and explicit representation of implicit temporal structures. During the online “wake phase,” efficient sequence learning is performed over this structured representation. Contribution/Results: This work introduces the first temporal chunking paradigm that jointly couples community detection with offline label generation, enabling cross-task knowledge transfer. On synthetic benchmarks, it significantly improves learning efficiency. Human behavioral experiments (Serial Reaction Time task) validate both the effectiveness of structural abstraction and the transferability of generated labels. The framework offers a novel, brain-inspired approach to sequence modeling and transfer learning under resource constraints.

Enhancing recognition of implicit sequential patterns using temporal chunkingExploring transfer learning potential via context tags across related tasksOvercoming limitations of RNNs in multi-timescale temporal patterns

Training Neural Networks as Recognizers of Formal Languages

Nov 11, 2024
AB
Alexandra Butoi
🏛️ ETH Zürich | Duke University | University of Copenhagen

Prior work commonly evaluates neural networks’ formal language recognition capabilities via proxy tasks like language modeling, creating a significant misalignment with formal language theory—which fundamentally concerns binary string classification. Method: We propose a theoretically grounded empirical paradigm: directly training RNNs, LSTMs, and causal Transformers as binary classifiers; designing a length-controllable regular language sampling algorithm (an improvement over Snæbjarnarson et al., 2024); and introducing FLaRe—the first benchmark dedicated to formal language recognition. Contribution/Results: Experiments reveal that RNNs and LSTMs consistently outperform causal Transformers across most languages in the Chomsky hierarchy; auxiliary objectives exhibit architecture- and language-specific efficacy. FLaRe is publicly released, establishing a new foundation for theoretically rigorous evaluation of AI’s reasoning capabilities.

Discrepancy between empirical tests and formal language theory claims.Neural networks' computational power in formal language recognition.Training neural networks as binary classifiers for string recognition.

Latest Papers

What's happening recently
View more

This work investigates whether large language models must rely on trainable input embedding tables. To address this, the authors propose a novel approach that entirely eliminates trainable input embeddings by employing fixed 16-dimensional binary token codes, combined with a zero-parameter dimensional expansion and an invertible affine recoding mechanism over a finite field, while retaining the standard trainable output projection. Evaluated on a 32-layer decoder model, this method achieves validation perplexity comparable to the baseline (2.36 vs. 2.44), reducing input parameters by approximately 67.1 million. A vocabulary-independent variant of the approach also performs closely (2.39). This study presents the first demonstration that large-scale language models can completely dispense with trainable input embeddings without sacrificing performance.

binary token codesembedding tableinput embedding

This work addresses the inflexibility of existing test-time training (TTT) methods, which are typically implemented as monolithic, hard-coded systems that hinder modular design and component-wise analysis. To overcome this limitation, the authors propose the first modular TTT framework, modeling the internal learner as a directed acyclic graph that explicitly decouples key elements such as fast weight networks, loss functions, and learning rates. The framework automatically composes elementary forward, backward, and query rules to construct complete computational pipelines, enabling systematic ablation studies and flexible reconfiguration. Through this approach, the study reveals the critical roles of small learning rate initialization, weight decay, and single-layer nonlinearity in achieving strong performance. Models built within this framework—scaled to 410 million and 1.45 billion parameters and trained on 100 billion tokens—match the training loss and downstream performance of Gated DeltaNet.

component analysisfast weightsmodular design

This work addresses the challenge of achieving both low-latency streaming synthesis and high audio quality in large language model (LLM)-based text-to-speech (TTS) systems, which typically rely on cumbersome GMM-HMM forced alignment tools like MFA. The authors propose CTC-TTS, the first framework to integrate a CTC-based neural aligner into LLM-based TTS, replacing the conventional MFA pipeline. A dual-token interleaving strategy is introduced to more accurately model text–speech alignment. Two variants are developed: CTC-TTS-L enhances audio quality through token concatenation, while CTC-TTS-F reduces latency via embedding stacking. Experimental results demonstrate that the proposed approach outperforms both MFA-based and fixed-interleaving baselines in streaming and zero-shot settings, significantly improving speech naturalness and response speed.

dual-streamingLLM-based TTSlow-latency

This work addresses the challenge of diacritic omission in Arabic speech transcription, which obscures fine-grained phonetic distinctions and hinders accurate modeling. The authors propose a CTC-based non-autoregressive approach for diacritic restoration that constructs a character-level diacritic lattice and enforces hard constraints during decoding to retain only linguistically valid diacritized sequences. By explicitly leveraging acoustic information and drastically reducing the decoding search space through these constraints, the method achieves superior accuracy while maintaining computational efficiency. Evaluated on the ArVoice and ClArTTS benchmarks, the proposed approach significantly lowers diacritic error rates compared to existing multimodal baselines, demonstrating both effectiveness and practicality for real-world deployment.

Arabic speech transcriptsdiacritic restorationphonological distinctions

This study addresses the limitation of conventional ASR training, which assumes unique transcriptions and overlooks local ambiguities. While existing Optimal Transducer Criterion (OTC) methods offer fault tolerance, they operate solely at the word level and risk discarding valid supervision. To overcome this, we propose a token-level OTC approach that refines wildcard arcs to subword granularity and integrates complementary word-level paths, thereby precisely preserving valid supervision signals. Furthermore, we introduce a predictive entropy-indexed scheduling mechanism that decouples length dependencies during training and dynamically localizes ambiguous regions. Extensive experiments demonstrate that the proposed method consistently outperforms CTC baselines across 25 tasks spanning 19 languages, achieving an average relative word error rate (WER) reduction of 9.45% and establishing state-of-the-art performance on all evaluated corpora.

Automatic Speech RecognitionConnectionist Temporal ClassificationOmni-temporal Classification

Hot Scholars

SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
XW

Xixin Wu

The Chinese University of Hong Kong
MC

Mingyu Cui

The Chinese University of Hong Kong
Speech RecognitionMachine Learning
SM

Shivam Mehta

Research Scientist Netflix - PhD @ KTH Royal Institute of Technology & WASP AI
Probabilistic Machine LearningDeep LearningSpeech SynthesisGenerative Models
MH

Michal Hradiš

Brno University of Technology
Computer VisionPattern Recognition