RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

πŸ“… 2025-05-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing cross-modal contrastive pretraining methods for the emerging task of Emotional Speaking Style Retrieval (ESSR) rely on strict binary audio–text alignment assumptions, limiting their ability to model complex emotional style relationships between speech and natural language descriptions. Method: We formally define ESSR and propose a relation-enhanced cross-modal pretraining framework. It introduces a self-distillation-based local matching learning mechanism to decouple global contrastive constraints, enabling fine-grained semantic associations (e.g., one-to-many, many-to-one). Combined with multi-granularity emotional feature modeling and CLAP architecture optimization, the framework achieves hierarchical alignment between speech and style descriptions. Contribution/Results: On the standard ESSR benchmark, our method significantly outperforms baselines including CLAP (+8.2% mAP), demonstrating that explicit relational modeling substantially improves generalization in cross-modal emotional understanding.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
πŸ“ Abstract
The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the emerging field of emotional speaking style description (ESSD), cross-modal contrastive pretraining remains largely unexplored. In this paper, we propose a novel speech retrieval task called emotional speaking style retrieval (ESSR), and ESS-CLAP, an emotional speaking style CLAP model tailored for learning relationship between speech and natural language descriptions. In addition, we further propose relation-augmented CLAP (RA-CLAP) to address the limitation of traditional methods that assume a strict binary relationship between caption and audio. The model leverages self-distillation to learn the potential local matching relationships between speech and descriptions, thereby enhancing generalization ability. The experimental results validate the effectiveness of RA-CLAP, providing valuable reference in ESSD.
Problem

Research questions and friction points this paper is trying to address.

Enhancing emotional speaking style retrieval via cross-modal contrastive learning
Addressing binary relationship limitations in caption-audio matching
Improving generalization with self-distilled local speech-description relationships
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emotional speaking style CLAP model
Relation-augmented CLAP for speech retrieval
Self-distillation for local matching relationships
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Haoqin Sun
Haoqin Sun
Nankai University
Affective computingSpeech signal processingAudio understanding
Jingguang Tian
Jingguang Tian
Midea AI Research Institute, Shanghai, China
speech LLMaudio understanding
J
Jiaming Zhou
TMCC, College of Computer Science, Nankai University, Tianjin, China
H
Hui Wang
TMCC, College of Computer Science, Nankai University, Tianjin, China
J
Jiabei He
TMCC, College of Computer Science, Nankai University, Tianjin, China
Shiwan Zhao
Shiwan Zhao
Independent Researcher, Research Scientist of IBM Research - China (2000-2020)
AGILarge Language ModelNLPSpeechRecommeder System
X
Xiangyu Kong
University of Exeter, Exeter, United Kingdom
D
Desheng Hu
Hithink RoyalFlush AI Research Institute, Zhejiang, China
X
Xinkang Xu
Hithink RoyalFlush AI Research Institute, Zhejiang, China
X
Xinhui Hu
Hithink RoyalFlush AI Research Institute, Zhejiang, China
Y
Yong Qin
TMCC, College of Computer Science, Nankai University, Tianjin, China