Score
Designs and implements algorithms and models that align unsegmented input sequences to target label sequences using Connectionist Temporal Classification (CTC), producing timestep- or frame-level alignments without requiring frame-level supervision or explicit lexicons. These solutions handle variable-length repeats, deletions, and noisy or time-warped inputs to predict normalized, aligned output sequences from raw sequential observations.
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
To address the difficulty of achieving accurate alignment in Connectionist Temporal Classification (CTC)-based automatic speech recognition (ASR) when speech and text sequences exhibit comparable lengths, this paper proposes an efficient end-to-end alignment framework. Our method introduces three key innovations: (1) a time-independent loss (TIL) that decouples frame-level prediction dependencies, enhancing robustness to misalignments; (2) an alignment-aware cross-entropy (AXE) loss explicitly modeling sequence matching via edit distance; and (3) a key-frame detection module coupled with a weighted fusion mechanism, reducing speech frames by over 86% on AISHELL-1/2 subsets while preserving discriminative acoustic information. Experiments demonstrate that our approach achieves significantly improved recognition accuracy over state-of-the-art CTC and alignment-enhanced baselines, without compromising low-latency inference.
CTC, while computationally efficient for automatic speech recognition (ASR), suffers from suboptimal performance due to its sharp output distribution and lack of contextual modeling. To address this, we propose Consistency-Regularized CTC (CR-CTC), the first method to introduce self-consistency regularization into the CTC framework. CR-CTC generates multiple augmented views of mel-spectrogram inputs, applies independent CTC modeling to each view, and enforces consistency among their output distributions via KL-divergence constraints. This yields implicit self-distillation and context-aware temporal masking representation learning, effectively mitigating CTC’s peaky output problem. Crucially, CR-CTC requires no architectural modifications or auxiliary decoders—only a principled loss-function redesign. Evaluated on LibriSpeech, AISHELL-1, and GigaSpeech, CR-CTC consistently outperforms standard CTC, matches or exceeds the accuracy of RNN-T and CTC/attention hybrid systems, and achieves state-of-the-art performance across benchmarks.
Multi-class variable-duration (MVD) segmented time series classification faces two key challenges: (1) neglect of temporal dependencies between adjacent segments and (2) inconsistent annotation boundaries. Method: Departing from the i.i.d. assumption, this work formally proves— for the first time—the discriminative gain conferred by contextual information. We propose a two-level context-prior-driven consistency learning framework integrating: (i) context-aware consistency regularization, (ii) dynamic neighborhood contrastive learning, (iii) soft boundary label smoothing, and (iv) temporal-adaptive feature alignment. Contribution/Results: These components jointly model inter-segment temporal dependencies and mitigate boundary annotation noise. Evaluated on multiple benchmark datasets, our method achieves average accuracy improvements of 3.2%–7.8%, demonstrating significantly enhanced robustness and tolerance to labeling noise.
This work proposes a multimodal generative framework that reformulates time series classification as a text generation task, jointly modeling numerical sequences, textual context, and task instructions. Traditional approaches often struggle to incorporate contextual information and overlook semantic relationships among classes. To address these limitations, the framework employs time series discretization, an alignment projection layer, and generative self-supervised pretraining, complemented by an implicit feature augmentation mechanism that integrates statistical features with vision-language image descriptions. This design effectively compensates for the inductive bias deficiencies of language models in temporal modeling. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method significantly outperforms existing approaches, highlighting its superior performance and strong generalization capability.
This work addresses the limited parallelizability of the classical dynamic time warping (DTW) algorithm, which suffers from quadratic time and memory complexity. The authors propose Segmental DTW, a novel approach that decomposes global sequence alignment into local subsequence DTW computations that can be executed in parallel, followed by a segment-level dynamic programming step to integrate the partial alignments. This method achieves near-full parallelism while preserving alignment accuracy comparable to standard DTW. Theoretical analysis and empirical evaluation on Chopin Mazurka audio alignment tasks demonstrate that one variant of the proposed method outperforms existing approaches in both computational efficiency and alignment performance.
Standard scaled dot-product attention lacks explicit modeling of continuous monotonic alignment, limiting performance on frame-synchronous tasks such as text-to-speech (TTS). To address this, we propose stochastic clock attention: it models source–target sequence alignment as the meeting probability of two learned non-negative stochastic clocks, and derives a closed-form Gaussian scoring function via path integral theory—ensuring causality, smoothness, and near-diagonal preference. This mechanism intrinsically enforces continuous monotonic alignment without positional regularization, supports both normalized and unnormalized forms, and unifies parallel and autoregressive decoding. In TTS, it significantly improves alignment stability and robustness to global temporal scaling variations, while maintaining or surpassing baseline models in speech quality.
This study addresses the absence of statistical learning theory for inference-time alignment under unknown rewards by formulating it as a weak-to-strong learning process. We introduce the concept of "alignment dimension" and prove that it fully characterizes learnability in this setting. By integrating the PAC learning framework with single-inclusion graph algorithms, we achieve effective alignment without reward estimation. This work establishes a comprehensive theoretical foundation for inference-time alignment in unknown reward scenarios, bridging critical gaps in learnability characterization and theoretical guarantees. Consequently, it provides a rigorous statistical basis for developing alignment algorithms, ensuring their reliability even when ground-truth rewards are inaccessible during inference.
This study addresses the limitation of conventional ASR training, which assumes unique transcriptions and overlooks local ambiguities. While existing Optimal Transducer Criterion (OTC) methods offer fault tolerance, they operate solely at the word level and risk discarding valid supervision. To overcome this, we propose a token-level OTC approach that refines wildcard arcs to subword granularity and integrates complementary word-level paths, thereby precisely preserving valid supervision signals. Furthermore, we introduce a predictive entropy-indexed scheduling mechanism that decouples length dependencies during training and dynamically localizes ambiguous regions. Extensive experiments demonstrate that the proposed method consistently outperforms CTC baselines across 25 tasks spanning 19 languages, achieving an average relative word error rate (WER) reduction of 9.45% and establishing state-of-the-art performance on all evaluated corpora.
This work addresses the challenge of achieving both low-latency streaming synthesis and high audio quality in large language model (LLM)-based text-to-speech (TTS) systems, which typically rely on cumbersome GMM-HMM forced alignment tools like MFA. The authors propose CTC-TTS, the first framework to integrate a CTC-based neural aligner into LLM-based TTS, replacing the conventional MFA pipeline. A dual-token interleaving strategy is introduced to more accurately model text–speech alignment. Two variants are developed: CTC-TTS-L enhances audio quality through token concatenation, while CTC-TTS-F reduces latency via embedding stacking. Experimental results demonstrate that the proposed approach outperforms both MFA-based and fixed-interleaving baselines in streaming and zero-shot settings, significantly improving speech naturalness and response speed.