assess sociolinguistic reasoning

Designs and implements evaluation frameworks and scoring systems that measure sociolinguistic reasoning and spoken emotional/empathetic intelligence (SpeechEQ/Spoken EQ) in multi‑turn spoken interactions. Builds analytic pipelines and annotation schemes to score paralinguistic and cross‑modal cues, diagnose modality shortcuts and safety traps, and validate reliability and fairness of spoken‑interaction assessments.

assesssociolinguisticreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited ability of existing spoken dialogue models to reason across modalities about paralinguistic social cues, which hinders their capacity to demonstrate authentic emotional intelligence (EI) in multi-turn conversations. To bridge this gap, the study introduces the first EI evaluation framework tailored for multi-turn spoken dialogues, grounded in the EQ-i 2.0 theoretical model. It presents a novel dataset comprising 2,265 multi-turn dialogues and proposes a Spoken EQ (SEQ) scoring protocol inspired by human EI assessment practices. Through multimodal speech-language modeling and paralinguistic feature analysis, the research identifies three critical bottlenecks in current systems: “modality shortcuts,” “safety traps,” and “contextual amnesia.” Experiments reveal that end-to-end spoken language models outperform cascaded architectures yet remain overly reliant on textual content. The dataset and an interactive demo platform are publicly released.

Emotional Intelligence QuotientMultimodal DialogueParalinguistic Cues

This work addresses the limitation of existing speech-language models in emotional intelligence evaluation, which often relies on shallow paralinguistic features without grounding in cognitive theory. Drawing upon the four-branch model of emotional intelligence, the study introduces EmoSBench—the first theory-driven benchmark—and proposes EmoS, a novel model optimized through supervised fine-tuning (SFT) and Grouped Relative Policy Optimization (GRPO). EmoS incorporates a joint reward mechanism combining Steep Exponential Accuracy Reward (SEAR) and Reasoning Fidelity Reward (RFR), trained on EmoDialogue, a newly curated fine-grained bilingual conversational dataset. Experimental results demonstrate that EmoS achieves 83.8% accuracy on EmoSBench, approaching human-level performance, and exhibits strong generalization capabilities in real-world, unconstrained spoken interactions.

Emotional IntelligenceEvaluation BenchmarkParalinguistic Perception

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

Oct 26, 2025
LZ
Li Zhou
🏛️ The Chinese University of Hong Kong, Shenzhen | Shenzhen Research Institute of Big Data

Current speech-language models (SLMs) lack joint perception and empathic reasoning over non-lexical acoustic cues—such as prosody, rhythm, and emotional intonation—while prevailing benchmarks evaluate isolated capabilities, failing to reflect the multimodal integration essential for authentic empathic dialogue. Method: We propose EchoMind, the first relational, multi-level benchmark that innovatively combines semantically neutral text with controllable voice styles to construct a fine-grained empathic evaluation framework covering 39 acoustic attributes, assessed across four interdependent stages: spoken language understanding, acoustic perception, holistic reasoning, and empathic response generation. Contribution/Results: Evaluating 12 state-of-the-art SLMs reveals consistent deficiencies in expressive speech recognition and context-adaptive empathic generation, particularly in instruction following and robustness to natural speech variability.

Assessing empathetic responses aligned with emotional and contextual factorsEvaluating SLMs' ability to perceive non-lexical vocal cues with spoken wordsTesting integration of linguistic, acoustic, reasoning and dialogue abilities

EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems

Aug 24, 2025
JL
Jingwen Liu
🏛️ Zhejiang University | UC Berkeley | University of Southern California | National Taiwan University

Current spoken dialogue systems lack systematic evaluation of affective reasoning capabilities—particularly cross-turn emotional coherence. To address this gap, we propose the first benchmark framework specifically designed for assessing emotional coherence in speech-based dialogues. Our method innovatively introduces a cross-turn affective reasoning scoring mechanism and leverages text-to-speech (TTS) synthesis to generate diverse, high-fidelity spoken evaluation data spanning multiple emotion categories and intensity levels. The framework integrates three complementary metric types: continuous (e.g., emotion intensity trajectory), categorical (e.g., polarity consistency), and perceptual (i.e., human subjective judgments), enabling multidimensional, reproducible assessment. Extensive experiments across seven state-of-the-art dialogue systems reveal prevalent patterns of emotional inconsistency, demonstrating the framework’s effectiveness and generalizability in detecting and quantifying emotional coherence deficits.

Addressing scarcity of emotional speech data for benchmarkingAssessing emotional coherence and transitions in multi-turn dialoguesLacking holistic evaluation for emotional reasoning in dialogue systems

This work addresses the shallow modeling of emotion understanding in spoken dialogue systems by proposing an Injective Emotion Attribution Thinking (IEAT) mechanism, which enables models to implicitly capture user emotions and their underlying causes during internal reasoning without relying on explicit supervision. Through a two-stage progressive training framework that integrates speech-text alignment, self-distillation, and cross-modal joint optimization, the approach achieves end-to-end modeling of emotional consistency and generation of empathetic responses. Evaluated on the HumDial benchmark, the method achieves state-of-the-art performance across three key tasks: emotional trajectory modeling, emotion attribution reasoning, and empathetic response generation, demonstrating superior results in both large language model–based and human evaluations.

emotion-aware reasoningemotional intelligenceempathetic response generation

Latest Papers

What's happening recently
View more

This study addresses the limited efficacy of existing dialogue systems in providing deep emotional support due to insufficient validation of user emotions. To bridge this gap, the authors introduce MEGUMI, the first systematic multilingual emotion validation framework, accompanied by the release of the M-EDESConv corpus and the M-TESC benchmark dataset. Built upon frozen XLM-RoBERTa semantic representations, MEGUMI integrates language-specific emotion encoders and employs cross-modal attention with gating mechanisms to jointly identify validating responses, detect optimal timing for validation, and generate empathetic replies. Experimental results demonstrate that MEGUMI significantly outperforms baseline models across multiple languages, achieving strong performance on both automatic metrics and human evaluations. The findings also highlight persistent limitations in current large language models’ capacity for nuanced emotion understanding.

dialogue systememotion understandingemotional validation

This study addresses the superficiality of empathetic interaction and the cascading propagation of reasoning errors in large speech models by reformulating empathetic dialogue as a structured cognitive process comprising perception, mental state reasoning, and response generation. Methodologically, this work proposes an audio-anchored attention mechanism to enhance acoustic representations, introduces step-decomposed credit assignment to precisely suppress error propagation, and employs a training paradigm combining supervised fine-tuning with reinforcement learning on a multi-stage empathy dataset. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance in perception, reasoning, and response alignment, significantly improving the interaction quality of deep emotional support.

Cognitive reasoningEmotional interactionEmpathetic spoken dialogue

This study addresses the challenge of precisely distinguishing turns, feedback, and pauses in spontaneous dialogue by proposing a semi-automatic annotation pipeline. The method integrates voice activity detection, energy filtering, automatic speech recognition, and contextual post-processing to automatically extract turns and feedback while generating consistent initial annotations to facilitate manual review. Experimental results demonstrate that the pipeline achieves an overall F1 score of 0.621 with a boundary error of approximately 0.15 seconds, exhibiting robustness to variations in listening conditions. By standardizing the dialogue annotation workflow, this work significantly enhances the reproducibility of conversational dynamics analysis.

backchannelsconversational turnsspeech-unit annotation

This study addresses the tendency of current large language models to conflate superficial politeness with deep emotional reasoning, often failing to distinguish between perceptual, cognitive, and interactive dimensions of emotional intelligence. Building upon the Mayer-Salovey-Caruso four-branch ability model, the authors propose FACET—a psychometric evaluation framework comprising 480 expert-designed items—to systematically assess model capabilities in emotion perception, facilitation, understanding, and management. The research reveals, for the first time, that emotional intelligence in large models is fragmented, identifying three distinct performance profiles: cognition-dominant, interaction-dominant, and context-dependent. A consistent bottleneck emerges in recognizing hidden emotions. While state-of-the-art models excel at emotion recognition and social reasoning, they struggle with effective interpersonal interaction, suggesting that current reinforcement learning from human feedback (RLHF) may induce only “stochastic empathy” rather than integrated emotional reasoning.

affective reasoningalignmentemotional intelligence

Hot Scholars

WZ

Wajdi Zaghouani

Associate Professor, Northwestern University
Digital HumanitiesComputational Social SciencesArabic Natural Language ProcessingComputational
HA

Héctor Allende-Cid

Fraunhofer IAIS / PUCV
Machine LearningData ScienceDistributed ComputingNatural Language Processing
JL

Jay L. Cunningham

DePaul University
Human Computer InteractionHuman-Centered ComputingResponsible AIInclusive Design
NO

Nedjma Ousidhoum

Lecturer (Assistant Professor), Cardiff University
Natural Language ProcessingComputational Social ScienceMachine Learning