Score
Design and implement training procedures that apply auxiliary Connectionist Temporal Classification (CTC) loss to intermediate encoder layers (inter-CTC) to encourage incremental alignment learning and stabilize encoder optimization. Build, weight, and schedule these intermediate losses and analyze their effects on encoder representations, convergence behavior, and sequence-level error metrics.
CTC-based ASR models offer efficient non-autoregressive decoding but suffer from weak language modeling capability, limiting recognition accuracy. To address this, we propose an LLM-driven intermediate loss regularization framework: leveraging the LLaMA embedding space as semantic supervision, we map intermediate representations from a Conformer encoder into the LLM’s semantic space via a learnable projection layer and jointly optimize with a causal language modeling loss. Crucially, this approach enhances speech features with linguistic knowledge without modifying the CTC decoding pipeline. Experiments on LibriSpeech, TEDLIUM2, and WSJ demonstrate substantial WER reductions, achieving state-of-the-art performance among CTC-based methods with negligible computational overhead. Our key contribution is the first use of an LLM’s semantic space as an intermediate supervision signal, enabling lightweight, efficient, and end-to-end language-aware regularization.
To address catastrophic forgetting—a critical bottleneck limiting both performance and efficiency in continual learning—this paper systematically identifies, for the first time, the strong forgetting-resistance of intermediate-layer representations. Building upon this insight, we propose a plug-and-play auxiliary classifier (AC) architecture that requires no modification to the backbone network or training pipeline. By embedding lightweight classifiers at intermediate layers and integrating an early-exit inference mechanism, our method jointly optimizes accuracy and efficiency. Under a multi-stage continual learning evaluation framework, it achieves an average relative accuracy gain of 10%, reduces inference computational cost by 10–60%, and preserves original accuracy—while remaining fully compatible with mainstream continual learning paradigms. Our core contributions are threefold: (i) uncovering the previously unrecognized forgetting-resistance of intermediate representations; (ii) introducing a modular, plug-and-play AC architecture; and (iii) establishing a new continual learning paradigm that simultaneously balances accuracy and efficiency.
CTC, while computationally efficient for automatic speech recognition (ASR), suffers from suboptimal performance due to its sharp output distribution and lack of contextual modeling. To address this, we propose Consistency-Regularized CTC (CR-CTC), the first method to introduce self-consistency regularization into the CTC framework. CR-CTC generates multiple augmented views of mel-spectrogram inputs, applies independent CTC modeling to each view, and enforces consistency among their output distributions via KL-divergence constraints. This yields implicit self-distillation and context-aware temporal masking representation learning, effectively mitigating CTC’s peaky output problem. Crucially, CR-CTC requires no architectural modifications or auxiliary decoders—only a principled loss-function redesign. Evaluated on LibriSpeech, AISHELL-1, and GigaSpeech, CR-CTC consistently outperforms standard CTC, matches or exceeds the accuracy of RNN-T and CTC/attention hybrid systems, and achieves state-of-the-art performance across benchmarks.
To address the challenges of continual expansion of novel categories, catastrophic forgetting mitigation, and elimination of reliance on large-scale fully supervised pretraining in few-shot class-incremental learning (FSCIL), this paper proposes Few-Shot Class-Incremental Tuning (FSCIT). FSCIT introduces a novel consistency-guided asynchronous contrastive tuning framework, integrating LoRA adaptation, dual-path asynchronous contrastive learning, controllable parameter freezing, and progressive consistency regularization—enabling efficient, low-forgetting category expansion without base-class pretraining. Evaluated across 16 benchmarks, FSCIT achieves an average 12.51% improvement over state-of-the-art methods, significantly alleviating forgetting and enhancing robustness under low-shot conditions. Under standard FSCIL settings, it yields an average gain of 2.47%, with up to +5.02% on individual datasets.
Neural processes often fail to accurately model dependencies between inputs and targets under noisy observations. To address this, we propose the Covariance Loss—a novel objective that explicitly incorporates second-order statistical dependencies among target variables into the end-to-end training of conditional neural processes for the first time. By regularizing the covariance structure of the predictive distribution, our loss enhances the model’s ability to recover missing or degraded dependencies and improves robustness to observation noise. The method is architecture-agnostic and can be seamlessly integrated into mainstream neural process frameworks. Extensive experiments across multiple real-world time-series and regression benchmarks demonstrate consistent and significant improvements over state-of-the-art methods in three key aspects: predictive accuracy, fidelity of dependency structure recovery, and robustness to observational noise.
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
This work addresses the inflexibility of existing test-time training (TTT) methods, which are typically implemented as monolithic, hard-coded systems that hinder modular design and component-wise analysis. To overcome this limitation, the authors propose the first modular TTT framework, modeling the internal learner as a directed acyclic graph that explicitly decouples key elements such as fast weight networks, loss functions, and learning rates. The framework automatically composes elementary forward, backward, and query rules to construct complete computational pipelines, enabling systematic ablation studies and flexible reconfiguration. Through this approach, the study reveals the critical roles of small learning rate initialization, weight decay, and single-layer nonlinearity in achieving strong performance. Models built within this framework—scaled to 410 million and 1.45 billion parameters and trained on 100 billion tokens—match the training loss and downstream performance of Gated DeltaNet.
Automatic phoneme recognition (APR) commonly relies on pseudo-phoneme labels generated by grapheme-to-phoneme (G2P) systems; however, standard Connectionist Temporal Classification (CTC) loss cannot model the inherent multi-pronunciation ambiguity in G2P outputs, resulting in poor robustness to label noise. To address this, we propose the first application of Graph-based Temporal Classification (GTC) to APR, wherein a phoneme sequence graph—explicitly encoding multiple pronunciation paths—serves as the supervision signal, enabling direct modeling of pronunciation uncertainty within the loss function. Our method supports end-to-end training without requiring post-processing or external pronunciation dictionaries. Evaluated on English and Dutch benchmark datasets, it achieves significant reductions in phoneme error rate, demonstrating both the effectiveness of integrating multi-pronunciation priors via GTC and its strong cross-lingual generalization capability.
This work addresses the challenge of insufficient accuracy and real-time performance in Arabic dialect identification under low-resource, streaming conditions. The authors propose a limited-vocabulary speech recognition framework based on Connectionist Temporal Classification (CTC) loss, which models dialect labels as sequences of phonetic units and enables end-to-end streaming inference. This study is the first to apply CTC to dialect identification and introduces a language-agnostic heuristic label repetition strategy that significantly enhances robustness for short utterances and zero-shot scenarios. By integrating self-supervised learning (SSL) models with CTC loss and leveraging alignment labels generated via LAH or pretrained ASR systems, the proposed approach outperforms fine-tuned Whisper and ECAPA-TDNN baselines on low-resource Arabic dialect recognition tasks, achieving superior performance particularly on the Casablanca dataset in zero-shot and short-duration evaluations.
This study addresses the limitation that sequential gradient writes during test-time training hinder parallel scaling. By revealing the duality between forward evaluation and backpropagation in online gradient descent, this work proposes a parallel scan algorithm based on costate prediction. Through the introduction of a causal auxiliary network and a consistency loss function, it enables exact parallel computation of forward inference and backpropagation alongside weight-efficient updates. The proposed mechanism strictly recovers the performance of sequential online learners. Furthermore, during deployment, simply removing the auxiliary network natively supports token-by-token model updates, offering a novel paradigm for efficient, parallelized test-time training.