🤖 AI Summary
This study addresses the high cost of human evaluation and the misalignment between automatic metrics and human judgment in speech emotion generation. To this end, it proposes SES-Bench, a dedicated benchmark, and SES-Judge, an evaluation model. Methodologically, this work introduces a novel human preference annotation framework based on pairwise comparisons and leverages large-scale listening test data to train a specialized neural network. This establishes an end-to-end paradigm for assessing emotional similarity that effectively supersedes conventional embedding cosine similarity and prompted large language model approaches. Experimental results demonstrate that the proposed model significantly outperforms existing baselines in preference accuracy and correlation with human ratings. By precisely capturing both the direction and magnitude of preferences, it achieves automated evaluation that closely aligns with human perception.
📝 Abstract
Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgments. To support the development and evaluation of such measures, we introduce SES-Bench, a speech emotion similarity benchmark built from human comparisons of two candidate utterances against a shared reference. These comparisons record which candidate listeners find emotionally closer to the reference and the strength of their preference. Using these annotations, we train SES-Judge to score emotion similarity between two utterances. SES-Judge significantly outperforms embedding cosine similarity and prompted large audio-language models in preference accuracy and correlation with human ratings that capture both preference direction and strength.