dyadic entrainment dataset creation

Designs and constructs datasets of paired (dyadic) interactions prepared for analysis of interpersonal entrainment, including procedures to record or compile conversational exchanges, segment them into analysis windows, and produce annotations indicating moments or degrees of entrainment. Builds controlled dataset variants and artifacts (e.g., partner-swapped or resynthesized interactions) to break or manipulate synchrony and documents the labeling and preprocessing needed for entrainment experiments.

dyadicentrainmentdatasetcreation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates emotional entrainment in dyadic spoken conversations, focusing on how social relationships and contextual dynamics shape affective coordination between interlocutors. To this end, the authors introduce DyadEE, a novel dataset comprising both authentic interactions and synthetically perturbed samples, and propose the TRACE framework. TRACE leverages Whisper acoustic embeddings fine-tuned for emotion recognition and models dialogues as window-level sequential interaction trajectories, incorporating relationship-aware mechanisms and temporal context modeling to capture dynamic entrainment patterns. Experimental results demonstrate that the proposed approach achieves a detection accuracy of 97.01% on DyadEE, underscoring the critical role of relational and contextual information in effectively modeling emotional entrainment in dyadic speech.

affective coordinationconversational contextdyadic speech

Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

Jun 27, 2025
VA
Vasu Agrawal
🏛️ Meta | University of Kansas

Modeling the dynamic multimodal coordination of speech, gestures, and facial expressions in face-to-face social interaction remains challenging for socially intelligent AI. Method: We introduce the first large-scale dyadic audio-visual interaction dataset (4,000+ hours) and propose a cross-modal sequential model integrating ASR, visual behavioral encoding, LLM-driven speech generation, and 2D/3D rendering to generate context-aware coordinated actions. A novel cross-modal alignment network enhances dyadic action prediction accuracy. Contribution/Results: Our framework enables fine-grained, emotion-state-, intensity-, and semantic-intent-conditioned controllable generation of gestures and facial expressions. Experiments demonstrate significant improvements in motion coherence and affective alignment of virtual agents. User studies confirm substantial gains in perceived naturalness and interaction quality, validating the efficacy of our approach for embodied social AI.

Create large-scale dataset for embodied dynamics analysisGenerate context-aware gestures and expressions from speechModel dyadic audiovisual behavior for human-AI interaction

Existing tools for dialogue research often lack modularity and adaptability, limiting their capacity to support diverse conversational contexts and experimental requirements. To address this gap, this work presents Dyadic (chatdyadic.com), a web-based, extensible platform that uniquely integrates multimodal text-and-speech interaction, real-time AI-generated suggestions, live researcher monitoring, and context-sensitive dynamic questionnaires within a unified system. Designed for zero-code configuration, Dyadic seamlessly interoperates with mainstream survey platforms, substantially enhancing experimental flexibility and ecological validity. By offering an out-of-the-box yet highly customizable environment for both human–human and human–agent dialogue studies, Dyadic lowers technical barriers and significantly expands the design space for empirical research in interactive communication.

conversation researchhuman-AI interactionhuman-human interaction

Multi-human Interactive Talking Dataset

Aug 04, 2025
ZZ
Zeyu Zhu
🏛️ National University of Singapore

Existing talking video generation research is largely confined to monologue scenarios or isolated facial animation, failing to model the bodily coordination and speech interaction inherent in realistic multi-person dialogues. To address this, we introduce MIT—the first large-scale dataset for multi-person interactive talking video generation—comprising 12 hours of high-resolution, naturally occurring dialogue videos with 2–4 participants, accompanied by fine-grained multi-body pose and speech interaction annotations. We further propose CovOG, a benchmark model designed to handle variable participant counts, featuring a Multi-Person Pose Encoder (MPE) and an Interactive Audio-Driven (IAD) module to explicitly model cross-speaker motion coupling and speech-responsive dynamics. Automated acquisition and annotation ensure high data fidelity. Experiments demonstrate that CovOG significantly improves motion naturalness and lip-sync accuracy over prior methods. Together, the MIT dataset and CovOG establish a new foundation for research in multi-person interactive talking video generation.

Addressing lack of datasets for multi-human talking video generationDeveloping automatic pipeline for multi-person conversational video annotationProposing baseline model for realistic multi-speaker interaction synthesis

为解决共处助手如何有效提供帮助的问题,通过构建同步多模态数据集DYAD,记录并分析人类在齿轮箱组装过程中请求与提供帮助的行为模式。

Co-located AssistanceHelp SeekingHuman-human Interaction

Latest Papers

What's happening recently
View more

This study addresses the latent biases inherent in diverse human-computer interaction datasets, which severely compromise the training and evaluation of user models. We introduce the concept of a "dataset signature" and employ neural network classifiers to perform source attribution across seven conversational datasets, conducting a systematic investigation integrating user modeling and preference analysis. Our results demonstrate that even under a unified taxonomy, individual datasets remain highly distinguishable, and variations in data sources can fundamentally alter conclusions regarding model quality. To address this, we propose a signature-based data filtering strategy. This approach offers a novel paradigm for mitigating dataset bias and enhancing the reliability of downstream tasks.

Data SelectionDataset BiasDataset Signatures

This work addresses the limitation of existing AI reflection tools in uncovering divergent interpretations between parties engaged in textual conflict. To bridge this gap, the authors propose a Dual-Spectator Reflection (DSR) mechanism that reframes conflicting individuals as co-observers of their interaction. The approach integrates AI-generated, revisable hypotheses about each participant’s internal states, pixel-art theater-style replays of key moments, a collaborative annotation interface, and “disagreement cards” to explicitly surface interpretive gaps at critical junctures. In an exploratory study with ten romantic couples, the system effectively enabled participants to transcend their individual perspectives and jointly revisit past conflicts through a shared, objective lens. This represents the first interactive paradigm specifically designed for collaborative reflection on relational conflict.

conversational divergencedyadic reflectioninterpretation gaps

Existing datasets struggle to support coupled analysis of affect across individual, interpersonal, and group levels in collaborative settings, and often suffer from fragmented, misaligned multimodal signals. To address this gap, this work introduces a high-ecological-validity multimodal dataset comprising synchronized physiological, eye-tracking, audio, continuous self-reported affect, personality traits, and task performance data from 10 four-person groups (40 participants total) engaged in four distinct collaborative tasks. All signals are temporally aligned and organized following a BIDS-inspired structure with Croissant metadata specifications. The dataset achieves 91% and 98% coverage for physiological and eye-tracking signals, respectively, and includes validation via affect manipulation checks. For the first time, it integrates a three-level analytical framework and provides leave-one-group-out cross-validation baselines alongside 15 reproducible benchmark tasks, offering a standardized, high-coverage resource for group affect research.

affective computingcollaborative tasksgroup interaction

This study addresses the fragmentation of findings in social and behavioral science experiments, which hinders the identification of knowledge consistency, contradictions, and research gaps. To tackle this issue, the authors propose ExAtlas—a novel framework that transforms archival social experiments into a structured, inferable knowledge graph for the first time. By leveraging the local smoothness assumption in treatment and outcome spaces, ExAtlas retrieves neighboring studies and aggregates their effects to predict outcomes of target experiments. The approach automatically links consistent evidence, explains apparent conflicts, and generates coherent bridging experiment proposals. Evaluated on target experiments with local support, the framework achieves 98.6% accuracy in predicting effect directions. Human assessments further confirm that the suggested bridging experiments are logically sound and that the conflict explanations meaningfully contribute to theory development.

evidence synthesisknowledge accumulationresearch conflicts