AMPS: ASR with Multimodal Paraphrase Supervision

📅 2024-11-27
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the poor automatic speech recognition (ASR) performance for low-resource Indian languages (e.g., Hindi, Marathi) and African languages (e.g., Chichewa) in spontaneous multilingual conversational ASR, this paper proposes Adaptive Multimodal Paraphrase Supervision (AMPS). AMPS introduces semantically equivalent paraphrases of reference transcripts as auxiliary supervision signals into multimodal ASR training—marking the first such integration—and employs an adaptive loss gating mechanism to dynamically activate paraphrase supervision for challenging recognition samples. Built upon the SeamlessM4T architecture, AMPS jointly models speech–text representation learning and paraphrase generation, enabling selective, semantics-aware optimization. Evaluated across five languages, AMPS achieves up to a 5% relative reduction in word error rate (WER), validated by both automated metrics and human evaluation. The method significantly enhances robustness and generalization of multilingual ASR in conversational settings.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLPPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics.
Problem

Research questions and friction points this paper is trying to address.

Improving multilingual conversational ASR accuracy
Using paraphrase supervision for poor-performance utterances
Reducing word error rates in multiple languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal ASR with paraphrase supervision
Selective paraphrase objective for poor utterances
Integration with SeamlessM4T for WER reduction
💼 Related Jobs
No related jobs found.
Indian Institute of Technology Bombay
A
Amruta Parulekar
Indian Institute of Technology Bombay, Mumbai, India
A
Abhishek Gupta
Indian Institute of Technology Bombay, Mumbai, India
Sameep Chattopadhyay
Sameep Chattopadhyay
Graduate Researcher, University of Washington
Machine LearningNLPSpeech RecognitionTime Series
P
P. Jyothi
Indian Institute of Technology Bombay, Mumbai, India