TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that current large language models struggle to capture the dynamics of long-term, multi-party group conversations in real-world settings, primarily due to a lack of high-resolution, longitudinal, naturally occurring multilingual dialogue data. To bridge this gap, we introduce the TIDES dataset, which tracks face-to-face meetings of 12 university teams over an entire semester, comprising 75,971 English–Korean bilingual utterances annotated for interaction types, emergent roles, and developmental stages. Leveraging this resource, we propose a fine-grained social structure annotation framework and fine-tune large language models for next-speaker prediction. Experimental results show that our TIDES-finetuned model achieves 64.53% accuracy—13.8 percentage points higher than a dyadic baseline—and approaches state-of-the-art performance on the AMI corpus while using 42% less training data, demonstrating its effectiveness in modeling authentic team dynamics.
📝 Abstract
Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challenge persists in part because existing group conversation datasets are often limited to short-term lab settings with contrived tasks, failing to capture the long-term social dynamics of real-world teams. To bridge this gap, we introduce TIDES, a high-resolution longitudinal dataset tracking 12 university project teams over a full semester. Comprising 75,971 utterances in both English and Korean from in-person meetings, TIDES provides a naturalistic record of teams working on self-managed projects. Our socio-structural annotations-covering interaction types, emergent roles, and development stages-allow for modeling of team evolution over months. Experiments show that fine-tuning on TIDES improves next-speaker prediction by 13.8 percentage points over a bigram baseline (64.53%) and yields performance comparable to strong proprietary zero-shot models. The model also comes within 2.1 percentage points of the published state of the art on the AMI Meeting Corpus while using approximately 42% less training data. However, human evaluations suggest that better next-speaker prediction does not necessarily yield more natural or coherent utterances, as fine-tuned models were generally less preferred than vanilla models. This potential mismatch motivates further study of how structural modeling can support natural multi-party generation.
Problem

Research questions and friction points this paper is trying to address.

multi-party conversation
social dynamics
longitudinal dataset
team interaction
next-speaker prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

longitudinal dataset
multi-party conversation
next-speaker prediction
socio-structural annotation
bilingual dialogue
🔎 Similar Papers
No similar papers found.