CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

📅 2024-11-12
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of poor speech-text modality alignment and weak cross-domain generalization in decoder-only end-to-end ASR, this paper proposes a CTC-compressor-driven joint training framework. Methodologically, it introduces (1) a novel bidirectional modality matching mechanism that integrates CTC compression, dynamic forced spike alignment, and CTC-inspired embeddings—enabling efficient audio-text fusion without explicit duration modeling; and (2) a lightweight modality adapter to enhance cross-domain robustness. Evaluated on LibriSpeech and TED-LIUM2, the approach achieves state-of-the-art performance among decoder-only models of comparable size, with significant improvements in noise robustness, long-form speech recognition, and cross-domain generalization. Furthermore, the study systematically characterizes optimal configurations of the CTC compressor under boundary and noisy conditions, providing principled guidance for its deployment in challenging acoustic scenarios.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Multimodal LearningCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, we propose a novel CTC compressor based joint speech and text training (CJST) framework for decoder-only ASR. CJST matches speech and text modalities from both directions by exploring a simple modality adaptor and several features of the CTC compressor, including sequence compression, on-the-fly forced peaky alignment and CTC class embeddings. Experimental results on the Librispeech and TED-LIUM2 corpora show that the proposed CJST achieves an effective text injection without the need of duration handling, leading to the best performance for both in-domain and cross-domain scenarios. We also provide a comprehensive study on CTC compressor, covering various compression modes, edge case handling and behavior under both clean and noisy data conditions, which reveals the most robust setting to use CTC compressor for decoder-only models.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition (ASR) Performance
Domain Adaptation
Speech and Text Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

CJST
CTC Compression
ASR Enhancement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Meta AI
W
Wei Zhou
Meta AI
J
Junteng Jia
Meta AI
L
Leda Sari
Meta AI
J
Jay Mahadeokar
Meta AI
O
Ozlem Kalinli
Meta AI