ctc-guided keyframe selection

Designs and implements algorithms that use Connectionist Temporal Classification (CTC) model outputs—peaky symbol/phoneme posterior frames—to select and extract high‑confidence temporal keyframes. Builds alignment modules that map audio frames to phoneme or text sequences to provide precise temporal alignment cues for multimodal processing.

ctc-guidedkeyframeselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

End-to-end Speech Recognition with similar length speech and text

Oct 12, 2025
PF
Peng Fan
🏛️ Chengdu University of Technology | Sichuan University

To address the difficulty of achieving accurate alignment in Connectionist Temporal Classification (CTC)-based automatic speech recognition (ASR) when speech and text sequences exhibit comparable lengths, this paper proposes an efficient end-to-end alignment framework. Our method introduces three key innovations: (1) a time-independent loss (TIL) that decouples frame-level prediction dependencies, enhancing robustness to misalignments; (2) an alignment-aware cross-entropy (AXE) loss explicitly modeling sequence matching via edit distance; and (3) a key-frame detection module coupled with a weighted fusion mechanism, reducing speech frames by over 86% on AISHELL-1/2 subsets while preserving discriminative acoustic information. Experiments demonstrate that our approach achieves significantly improved recognition accuracy over state-of-the-art CTC and alignment-enhanced baselines, without compromising low-latency inference.

Addressing speech-text length mismatch in ASR systemsEnhancing keyframe information through frame fusion techniquesImproving alignment accuracy with novel loss functions

This work addresses the challenge of distinguishing target keywords from phoneme-level confusables in user-defined keyword spotting. To this end, the authors propose a multimodal detection framework that leverages the peak characteristics of CTC posterior distributions to precisely select high-confidence keyframes, thereby enabling effective alignment across audio, phoneme, and text modalities. A cross-attention mechanism is further introduced to jointly exploit the local discriminability of keyframes and the global contextual information of the entire utterance. Evaluated on the LibriPhrase dataset, the proposed model achieves state-of-the-art performance with an AUC of 98.73% overall, and notably attains 97.65% AUC and 7.75% EER on the challenging subset, significantly outperforming existing approaches.

confusable keywordskeyword spottingmultimodal alignment

CTC, while computationally efficient for automatic speech recognition (ASR), suffers from suboptimal performance due to its sharp output distribution and lack of contextual modeling. To address this, we propose Consistency-Regularized CTC (CR-CTC), the first method to introduce self-consistency regularization into the CTC framework. CR-CTC generates multiple augmented views of mel-spectrogram inputs, applies independent CTC modeling to each view, and enforces consistency among their output distributions via KL-divergence constraints. This yields implicit self-distillation and context-aware temporal masking representation learning, effectively mitigating CTC’s peaky output problem. Crucially, CR-CTC requires no architectural modifications or auxiliary decoders—only a principled loss-function redesign. Evaluated on LibriSpeech, AISHELL-1, and GigaSpeech, CR-CTC consistently outperforms standard CTC, matches or exceeds the accuracy of RNN-T and CTC/attention hybrid systems, and achieves state-of-the-art performance across benchmarks.

Enhances contextual representation learningImproves CTC speech recognitionReduces overfitting in CTC distributions

To address the challenges of poor speech-text modality alignment and weak cross-domain generalization in decoder-only end-to-end ASR, this paper proposes a CTC-compressor-driven joint training framework. Methodologically, it introduces (1) a novel bidirectional modality matching mechanism that integrates CTC compression, dynamic forced spike alignment, and CTC-inspired embeddings—enabling efficient audio-text fusion without explicit duration modeling; and (2) a lightweight modality adapter to enhance cross-domain robustness. Evaluated on LibriSpeech and TED-LIUM2, the approach achieves state-of-the-art performance among decoder-only models of comparable size, with significant improvements in noise robustness, long-form speech recognition, and cross-domain generalization. Furthermore, the study systematically characterizes optimal configurations of the CTC compressor under boundary and noisy conditions, providing principled guidance for its deployment in challenging acoustic scenarios.

Automatic Speech Recognition (ASR) PerformanceDomain AdaptationSpeech and Text Integration

This work addresses the challenges of synchronous audio-video generation and weak cross-modal alignment in diffusion-based generative models. To this end, we propose a strong baseline framework for joint audio-video generation. Our method integrates pre-trained audio and video diffusion models while eliminating inefficient cross-attention mechanisms. Instead, we introduce two novel components: a timestep-aware module and cross-modal conditional positional encoding (CMC-PE), which explicitly encode temporal alignment priors. Furthermore, we employ multimodal feature fusion and end-to-end joint training. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across multiple benchmarks, significantly outperforming existing methods in three key dimensions—generation quality, temporal consistency, and cross-modal alignment. These results validate the effectiveness and generalizability of our designed inductive biases.

Enhance alignment between audio-video pairs with novel mechanismsImprove temporal alignment in generated data using CMC-PEIntegrate audio and video diffusion models for joint generation

Latest Papers

What's happening recently
View more

This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.

automatic speech recognitionencoder frame gridspeech LLMs

This work addresses the challenge of achieving both low-latency streaming synthesis and high audio quality in large language model (LLM)-based text-to-speech (TTS) systems, which typically rely on cumbersome GMM-HMM forced alignment tools like MFA. The authors propose CTC-TTS, the first framework to integrate a CTC-based neural aligner into LLM-based TTS, replacing the conventional MFA pipeline. A dual-token interleaving strategy is introduced to more accurately model text–speech alignment. Two variants are developed: CTC-TTS-L enhances audio quality through token concatenation, while CTC-TTS-F reduces latency via embedding stacking. Experimental results demonstrate that the proposed approach outperforms both MFA-based and fixed-interleaving baselines in streaming and zero-shot settings, significantly improving speech naturalness and response speed.

dual-streamingLLM-based TTSlow-latency

This work addresses the challenge of diacritic omission in Arabic speech transcription, which obscures fine-grained phonetic distinctions and hinders accurate modeling. The authors propose a CTC-based non-autoregressive approach for diacritic restoration that constructs a character-level diacritic lattice and enforces hard constraints during decoding to retain only linguistically valid diacritized sequences. By explicitly leveraging acoustic information and drastically reducing the decoding search space through these constraints, the method achieves superior accuracy while maintaining computational efficiency. Evaluated on the ArVoice and ClArTTS benchmarks, the proposed approach significantly lowers diacritic error rates compared to existing multimodal baselines, demonstrating both effectiveness and practicality for real-world deployment.

Arabic speech transcriptsdiacritic restorationphonological distinctions

This work addresses the limitations of existing audio-visual captioning methods, which often fail to accurately associate auditory events with visual entities or model complex causal dynamics due to modality misalignment and temporal inconsistency. To overcome these challenges, the authors propose TCA-Captioner, a novel framework featuring an Observer-Checker-Corrector (OCC) iterative refinement mechanism. It leverages high-density human-machine interaction data to generate high-fidelity training samples and employs a multimodal large language model to enhance cross-modal alignment. Additionally, the study introduces TCA-Bench, a diagnostic benchmark with a disentangled evaluation protocol that separately quantifies a model’s capabilities in audio-visual binding and temporal reasoning. Experiments demonstrate that the proposed approach significantly improves temporal coherence and cross-modal synchronization in generated captions, establishing a new state-of-the-art on TCA-Bench.

Audiovisual Video CaptioningCross-Modal AlignmentModality Detachment

This work addresses the susceptibility of unified audio-language models to temporal smoothing bias during generation, which hinders their effective utilization of transient acoustic cues and results in insufficient fine-grained alignment between output text and audio. To mitigate this issue, the authors propose a training-free temporal contrastive decoding method that, at inference time, constructs a contrastive signal between the original input and a temporally blurred “slow-path” view to dynamically refine the logits of the next token. The approach introduces, for the first time, a self-normalized stability score coupled with an uncertainty-aware gating mechanism, integrating waveform-smoothed recoding, adaptive blurring windows, and token-level logit updates to enable on-demand, precise enhancement of transient audio information. Evaluated on the MMAU and AIR-Bench benchmarks, the method consistently improves performance across multiple strong baseline models, demonstrating both effectiveness and architectural generality.

audio-grounded generationlarge audio-language modelstemporal smoothing bias