🤖 AI Summary
This study addresses the unclear capacity of existing speech-to-speech models to perceive and adaptively respond to non-verbal vocalizations (NSVs). To investigate this, we introduce a novel NSV evaluation framework based on contrastive pairing. Specifically, we construct a benchmark by embedding distinct NSVs into lexically identical dialogues, complemented by a human-validated data curation pipeline and a multidimensional assessment protocol that systematically evaluates model capabilities in NSV detection, affective comprehension, and response modulation. Experimental results reveal that current models perform substantially better at detecting NSVs than at interpreting fine-grained emotional nuances. Furthermore, we release the complete dataset and source code as open-source resources. This work establishes a standardized evaluation paradigm for processing non-verbal information in spoken dialogue systems, offering a valuable foundation for future research in expressive speech interaction.
📝 Abstract
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.