detect dyadic entrainment

Design and implement algorithms and analysis pipelines that quantify and detect entrainment between two interacting participants by measuring cross-speaker feature similarity and changes in temporal alignment. Build metrics and classifiers that compute turn-level synchrony (e.g., prosodic, lexical, or emotional alignment) and decide whether conversational or emotional entrainment is present.

detectdyadicentrainment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates emotional entrainment in dyadic spoken conversations, focusing on how social relationships and contextual dynamics shape affective coordination between interlocutors. To this end, the authors introduce DyadEE, a novel dataset comprising both authentic interactions and synthetically perturbed samples, and propose the TRACE framework. TRACE leverages Whisper acoustic embeddings fine-tuned for emotion recognition and models dialogues as window-level sequential interaction trajectories, incorporating relationship-aware mechanisms and temporal context modeling to capture dynamic entrainment patterns. Experimental results demonstrate that the proposed approach achieves a detection accuracy of 97.01% on DyadEE, underscoring the critical role of relational and contextual information in effectively modeling emotional entrainment in dyadic speech.

affective coordinationconversational contextdyadic speech

This study addresses the challenge of precisely distinguishing turns, feedback, and pauses in spontaneous dialogue by proposing a semi-automatic annotation pipeline. The method integrates voice activity detection, energy filtering, automatic speech recognition, and contextual post-processing to automatically extract turns and feedback while generating consistent initial annotations to facilitate manual review. Experimental results demonstrate that the pipeline achieves an overall F1 score of 0.621 with a boundary error of approximately 0.15 seconds, exhibiting robustness to variations in listening conditions. By standardizing the dialogue annotation workflow, this work significantly enhances the reproducibility of conversational dynamics analysis.

backchannelsconversational turnsspeech-unit annotation

This study investigates the spatiotemporal emotional synchronization between facial and vocal modalities during dyadic conversations, specifically along arousal and valence dimensions, with a focus on how speech overlap versus non-overlap conditions modulate cross-modal alignment. We employ EmoNet for continuous facial affect estimation and fine-tuned Wav2Vec2 for speech-based affect modeling, quantifying synchronization via Pearson correlation, lag analysis, and dynamic time warping (DTW). Our key findings reveal that speech overlap critically regulates synchronization stability and temporal lag patterns: non-overlapping segments exhibit higher synchronization stability (reduced arousal variability; narrower lag distributions), whereas overlapping segments—despite lower DTW distances—show flattened, highly uncertain lag profiles. Furthermore, we identify a structural shift in modality dominance: facial expressions lead turn-taking speech, while vocal signals lead overlapping speech. These results establish a novel “dialogue-structure-driven emotional coordination” paradigm.

Analyzing spatial-temporal dynamics of multimodal affective coordination in real-world interactionsExamining emotional synchrony in dyadic interactions across facial and vocal modalitiesInvestigating how speech overlap affects arousal and valence alignment in conversations

This study investigates differences between human interlocutors and classification models in their sensitivity to lexical, prosodic, and code-switching style cues during bilingual conversations, addressing a gap in cross-linguistic alignment research. Analyzing Mandarin–English, Hindi–English, and Spanish–English dyadic dialogues, the work combines feature importance analysis and ablation studies to compare how traditional classifiers and Transformer-based models respond to these multimodal signals. The research introduces the first cross-linguistic framework for modeling human alignment behavior and proposes a novel paradigm for evaluating multilingual dialogue systems against human behavioral benchmarks. Findings reveal that lexical alignment exhibits cross-linguistic universality, whereas prosodic and code-switching style alignment show language-pair specificity. Although models can detect alignment patterns, they rely on different feature cues than humans do.

code-switchingconversational entrainmentcross-lingual analysis

From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines

Dec 12, 2025
TM
Titaya Mairittha
🏛️ AXONS | Chulalongkorn University

This study identifies three structural dialogue fractures—temporal misalignment, expressive flattening, and rigid repair—in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) systems, arising from excessive component-level controllability at the expense of conversational fluidity. Using multimodal interaction experiments on production-grade systems, we integrate conversation analysis, latency-aware behavioral modeling, and architectural decoupling assessment. We establish, for the first time, that dialogue friction stems not from isolated technical flaws but from interface design mismatches across modules. Accordingly, we propose “interface orchestration” as a new paradigm—replacing conventional “component optimization”—to reframe infrastructure challenges in natural-speech AI. Our work formalizes reproducible fracture patterns and articulates architectural design principles that jointly ensure controllability and interactional fluency, providing theoretical foundations and practical guidance for next-generation spoken-language AI systems.

Analyzes temporal misalignment, expressive flattening, and repair rigidity as friction pointsIdentifies conversational breakdown patterns in modular speech-to-speech AI systemsProposes shifting from component optimization to choreographing system seams for natural interaction

Latest Papers

What's happening recently
View more

This study addresses the challenge of reliably measuring conversational states—such as cognitive load and conversational dominance—from multimodal behavioral signals, balancing predictive power, cross-task generalizability, and test–retest reliability. Leveraging the AVCAffe dataset (53 dyads across nine remote collaborative tasks), the authors construct a three-dimensional assessment framework integrating interactional, acoustic, and linguistic features, enhanced by speaker normalization to improve comparability. Findings reveal that while linguistic features exhibit the strongest predictive performance for cognitive load, they generalize poorly across tasks; acoustic features show high reliability but are strongly speaker-dependent; only interactional features—such as speaking dominance duration—robustly capture within-dyad asymmetries in cognitive load. Notably, classification of conversational dominance roles remains near chance-level across conditions. This work provides both methodological guidance and empirical grounding for selecting reliable multimodal features in social signal processing.

conversational statecross-task generalizabilitymultimodal signals

This study addresses the limited understanding of articulatory coordination mechanisms in spontaneous conversation and the constraints of invasive measurement techniques by proposing a speaker-independent, non-invasive analytical framework. Leveraging acoustic-to-articulatory inversion, this framework provides the first quantification of articulatory coordination complexity and synchrony in spontaneous dialogue. Experimental results demonstrate that neurotypical individuals exhibit significantly higher articulatory synchrony than autistic individuals, and that elevated synchrony positively correlates with subjective conversational success. By revealing dynamic interactional differences across neurotypes, this work establishes articulatory coordination as a valuable objective metric for social effectiveness.

Articulatory EntrainmentAutismCoordination Complexity

This study addresses the limitation of existing full-duplex spoken dialogue model evaluations, which rely on unidirectional interactions and fail to capture coupled behaviors between interlocutors. Adopting a novel "two-body problem" perspective, this work proposes DyaFDB, a framework enabling two models to engage in direct dialogue under assigned roles while performing bidirectional scoring, thereby establishing a closed-loop evaluation paradigm where each model serves as both examiner and examinee. By integrating full-duplex speech modeling, multi-agent game dynamics, and an offline automated judge system, the approach effectively simulates authentic interactive dynamics. Based on 7,560 dialogue sessions, the study reveals how interlocutor behaviors mutually reshape one another during interaction. The evaluation protocol is open-sourced, providing a new benchmark for the rigorous scientific assessment of full-duplex spoken dialogue systems.

Cross-playDyadic evaluationDyaFDB

Hot Scholars

PG

Pedro Guillermo Feijóo-García

Lecturer, School of Computing Instruction, Georgia Institute of Technology
Human-Centered ComputingVirtual HumansCS Education
ZE

Zag ElSayed

UC, CCHMC, OCR
Computer Engineering BCIEEGCybersecurityAI and ML