Toward Human-Aligned Judgement of Speech Emotion Similarity

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost of human evaluation and the misalignment between automatic metrics and human judgment in speech emotion generation. To this end, it proposes SES-Bench, a dedicated benchmark, and SES-Judge, an evaluation model. Methodologically, this work introduces a novel human preference annotation framework based on pairwise comparisons and leverages large-scale listening test data to train a specialized neural network. This establishes an end-to-end paradigm for assessing emotional similarity that effectively supersedes conventional embedding cosine similarity and prompted large language model approaches. Experimental results demonstrate that the proposed model significantly outperforms existing baselines in preference accuracy and correlation with human ratings. By precisely capturing both the direction and magnitude of preferences, it achieves automated evaluation that closely aligns with human perception.
📝 Abstract
Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgments. To support the development and evaluation of such measures, we introduce SES-Bench, a speech emotion similarity benchmark built from human comparisons of two candidate utterances against a shared reference. These comparisons record which candidate listeners find emotionally closer to the reference and the strength of their preference. Using these annotations, we train SES-Judge to score emotion similarity between two utterances. SES-Judge significantly outperforms embedding cosine similarity and prompted large audio-language models in preference accuracy and correlation with human ratings that capture both preference direction and strength.
Problem

Research questions and friction points this paper is trying to address.

Speech emotion similarity
Expressive speech generation
Human-aligned evaluation
Automatic metrics
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Emotion Similarity
Human-Aligned Evaluation
SES-Bench
SES-Judge
Expressive Speech Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yun-Shao Tsai
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan
Yi-Cheng Lin
Yi-Cheng Lin
National Taiwan University
Speech ProcessingMachine LearningFairness
Chih-Kai Yang
Chih-Kai Yang
National Taiwan University
Deep LearningSpeech ProcessingNatural Language ProcessingMachine Learning
H
Ho-Jung Cheng
Independent Researcher
T
Tsun-Yi Chang
National Taiwan University, Taiwan
S
Sheng-Wei Wu
National Taiwan University of Science and Technology, Taiwan
Y
Yi-Shan Chen
National Taiwan University of Science and Technology, Taiwan
H
Hsiang-Chun Chang
National Taiwan University, Taiwan
L
Liang-Chieh Lee
National Taiwan University, Taiwan
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing