Score
Designs and implements sequence models trained with Connectionist Temporal Classification (CTC) loss to map variable-length input sequences to output token sequences without frame-level alignment; this includes computing the CTC objective, optimizing model parameters, handling blank symbols and repeated-label collapsing, batching and numerical stability, and building decoding/beam-search and sequence-level evaluation components for alignment-free training.
CTC, while computationally efficient for automatic speech recognition (ASR), suffers from suboptimal performance due to its sharp output distribution and lack of contextual modeling. To address this, we propose Consistency-Regularized CTC (CR-CTC), the first method to introduce self-consistency regularization into the CTC framework. CR-CTC generates multiple augmented views of mel-spectrogram inputs, applies independent CTC modeling to each view, and enforces consistency among their output distributions via KL-divergence constraints. This yields implicit self-distillation and context-aware temporal masking representation learning, effectively mitigating CTC’s peaky output problem. Crucially, CR-CTC requires no architectural modifications or auxiliary decoders—only a principled loss-function redesign. Evaluated on LibriSpeech, AISHELL-1, and GigaSpeech, CR-CTC consistently outperforms standard CTC, matches or exceeds the accuracy of RNN-T and CTC/attention hybrid systems, and achieves state-of-the-art performance across benchmarks.
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
Connectionist Temporal Classification (CTC) theoretically assumes conditional independence among output labels, yet empirical evidence suggests that strong encoders implicitly learn context-dependent internal language models (ILMs). Method: This work formally models and empirically validates this implicit contextual dependency for the first time; proposes a label-level knowledge distillation framework to unsupervisedly extract context-aware ILMs from CTC decoders; and introduces two novel strategies—smooth regularization and label-level contextual modeling—to relax the conventional independent-label assumption. Contribution/Results: On the cross-domain TED-LIUM benchmark, the proposed approach reduces word error rate (WER) by over 13% compared to shallow fusion, significantly outperforming context-agnostic priors. These results robustly demonstrate CTC’s capacity to learn strong implicit ILMs, challenging the standard independence assumption and enabling more effective integration of contextual knowledge in end-to-end speech recognition.
Traditional RNNs struggle to capture implicit temporal patterns spanning multiple time scales. Method: Inspired by neuroscience, we propose a “wake–sleep” two-phase learning framework. During the offline “sleep phase,” a graph-community-detection–based temporal chunking mechanism automatically generates context-aware, structured memory units—enabling long-sequence compression and explicit representation of implicit temporal structures. During the online “wake phase,” efficient sequence learning is performed over this structured representation. Contribution/Results: This work introduces the first temporal chunking paradigm that jointly couples community detection with offline label generation, enabling cross-task knowledge transfer. On synthetic benchmarks, it significantly improves learning efficiency. Human behavioral experiments (Serial Reaction Time task) validate both the effectiveness of structural abstraction and the transferability of generated labels. The framework offers a novel, brain-inspired approach to sequence modeling and transfer learning under resource constraints.
Prior work commonly evaluates neural networks’ formal language recognition capabilities via proxy tasks like language modeling, creating a significant misalignment with formal language theory—which fundamentally concerns binary string classification. Method: We propose a theoretically grounded empirical paradigm: directly training RNNs, LSTMs, and causal Transformers as binary classifiers; designing a length-controllable regular language sampling algorithm (an improvement over Snæbjarnarson et al., 2024); and introducing FLaRe—the first benchmark dedicated to formal language recognition. Contribution/Results: Experiments reveal that RNNs and LSTMs consistently outperform causal Transformers across most languages in the Chomsky hierarchy; auxiliary objectives exhibit architecture- and language-specific efficacy. FLaRe is publicly released, establishing a new foundation for theoretically rigorous evaluation of AI’s reasoning capabilities.
This work investigates whether large language models must rely on trainable input embedding tables. To address this, the authors propose a novel approach that entirely eliminates trainable input embeddings by employing fixed 16-dimensional binary token codes, combined with a zero-parameter dimensional expansion and an invertible affine recoding mechanism over a finite field, while retaining the standard trainable output projection. Evaluated on a 32-layer decoder model, this method achieves validation perplexity comparable to the baseline (2.36 vs. 2.44), reducing input parameters by approximately 67.1 million. A vocabulary-independent variant of the approach also performs closely (2.39). This study presents the first demonstration that large-scale language models can completely dispense with trainable input embeddings without sacrificing performance.
This work addresses the inflexibility of existing test-time training (TTT) methods, which are typically implemented as monolithic, hard-coded systems that hinder modular design and component-wise analysis. To overcome this limitation, the authors propose the first modular TTT framework, modeling the internal learner as a directed acyclic graph that explicitly decouples key elements such as fast weight networks, loss functions, and learning rates. The framework automatically composes elementary forward, backward, and query rules to construct complete computational pipelines, enabling systematic ablation studies and flexible reconfiguration. Through this approach, the study reveals the critical roles of small learning rate initialization, weight decay, and single-layer nonlinearity in achieving strong performance. Models built within this framework—scaled to 410 million and 1.45 billion parameters and trained on 100 billion tokens—match the training loss and downstream performance of Gated DeltaNet.
This work addresses the challenge of achieving both low-latency streaming synthesis and high audio quality in large language model (LLM)-based text-to-speech (TTS) systems, which typically rely on cumbersome GMM-HMM forced alignment tools like MFA. The authors propose CTC-TTS, the first framework to integrate a CTC-based neural aligner into LLM-based TTS, replacing the conventional MFA pipeline. A dual-token interleaving strategy is introduced to more accurately model text–speech alignment. Two variants are developed: CTC-TTS-L enhances audio quality through token concatenation, while CTC-TTS-F reduces latency via embedding stacking. Experimental results demonstrate that the proposed approach outperforms both MFA-based and fixed-interleaving baselines in streaming and zero-shot settings, significantly improving speech naturalness and response speed.
This work addresses the challenge of diacritic omission in Arabic speech transcription, which obscures fine-grained phonetic distinctions and hinders accurate modeling. The authors propose a CTC-based non-autoregressive approach for diacritic restoration that constructs a character-level diacritic lattice and enforces hard constraints during decoding to retain only linguistically valid diacritized sequences. By explicitly leveraging acoustic information and drastically reducing the decoding search space through these constraints, the method achieves superior accuracy while maintaining computational efficiency. Evaluated on the ArVoice and ClArTTS benchmarks, the proposed approach significantly lowers diacritic error rates compared to existing multimodal baselines, demonstrating both effectiveness and practicality for real-world deployment.
This study addresses the limitation of conventional ASR training, which assumes unique transcriptions and overlooks local ambiguities. While existing Optimal Transducer Criterion (OTC) methods offer fault tolerance, they operate solely at the word level and risk discarding valid supervision. To overcome this, we propose a token-level OTC approach that refines wildcard arcs to subword granularity and integrates complementary word-level paths, thereby precisely preserving valid supervision signals. Furthermore, we introduce a predictive entropy-indexed scheduling mechanism that decouples length dependencies during training and dynamically localizes ambiguous regions. Extensive experiments demonstrate that the proposed method consistently outperforms CTC baselines across 25 tasks spanning 19 languages, achieving an average relative word error rate (WER) reduction of 9.45% and establishing state-of-the-art performance on all evaluated corpora.