Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of objective benchmarks and the alignment challenges associated with subjective dimensions, such as emotion and interpretation, in speech aesthetics. To this end, we construct a conversational large model for speech aesthetic evaluation. Methodologically, the approach combines human-annotated supervised fine-tuning (SFT) with a proposed Group Relative Policy Optimization (GRPO) framework, leveraging listener group feedback to achieve principled alignment of subjective aesthetic perception. Experimental results demonstrate that the proposed model surpasses Gemini 3.1 Pro and mainstream open-source baselines in agreement with human judgments, further exceeding the consistency level of single-annotator evaluations. This work establishes a novel paradigm for subjective speech aesthetic assessment.
📝 Abstract
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Voice Aesthetic Model
Reinforcement Learning from Human Feedback
Group Relative Policy Optimization
Speech Large Language Model
Human Alignment
🔎 Similar Papers
No similar papers found.