Score
Designs, builds, and evaluates algorithms, software, and datasets that detect, extract, fuse, and model human social signals (e.g., facial expressions, gaze, gesture, prosody, posture, turn‑taking) from multimodal interaction data. Uses those components to infer, track, or synthesize interpersonal states, intentions, engagement, rapport, and other social dynamics for real‑time or offline analysis.
This study addresses the scarcity of publicly available datasets that simultaneously capture multimodal behavioral signals and self-reported annotations, which has hindered the modeling of stance (agree/disagree/neutral) and social cues in dyadic conversations. To bridge this gap, the authors present a novel multimodal interaction corpus comprising 45 dyads (90 participants), encompassing both familiar and unfamiliar pairs. The dataset integrates synchronized recordings of 2D/3D facial video, thermal imaging, audio, and multiple physiological signals—including photoplethysmography (PPG), electrodermal activity (EDA), heart rate, blood pressure, and respiration—alongside structured stance labels and subjective emotional reports. Spanning approximately 20 terabytes, this resource is the first to jointly provide fine-grained behavioral data and introspective self-assessments, thereby enabling robust computational modeling of interpersonal social dynamics and filling a critical void in the field.
Existing research often examines textual, visual, or audio modalities of short videos in isolation, failing to uncover how their interplay influences user engagement. This work proposes the first reproducible and interpretable multimodal analysis framework that integrates automated feature extraction with Shapley-value-based attribution to systematically investigate how multimodal interactions affect view counts in TikTok content related to social anxiety disorder. The study reveals that facial expressions are more predictive than textual sentiment, that informational content garners greater attention than emotional support, and that multimodal synergies exhibit strong threshold-dependent effects—thereby transcending the limitations of conventional unimodal analyses.
In psychology and social sciences, manual annotation of multimodal behavioral data is costly, time-consuming, and suffers from low inter-annotator consistency. To address this, we present an open-source, modular, task-oriented multimodal behavioral feature extraction toolkit. The toolkit integrates speaker diarization, automatic speech recognition (ASR), machine translation, and vision-based pose estimation (e.g., MediaPipe and OpenPose) into an end-to-end configurable pipeline, enabling automated extraction of speaker identity, transcribed and translated speech, body/gesture/face pose, and gaze trajectories from audiovisual recordings. Its novel extensible architecture significantly lowers the barrier to entry for non-AI-expert researchers, promoting standardization and democratization of behavioral analysis tools. Empirical evaluation demonstrates high accuracy and plug-and-play usability, substantially improving annotation efficiency and consistency. The toolkit has already enabled multiple real-time studies on dynamic human behavior analysis.
This paper identifies four critical limitations in current LLM-driven multimodal human behavior understanding systems: (1) overreliance on the “modality-to-text” paradigm, neglecting fine-grained audiovisual social cues; (2) absence of adaptive interactive reasoning capabilities; (3) evaluation confined to static benchmarks, lacking social context and human-centered perspectives; and (4) ethical discourse focused predominantly on legal risks while overlooking socially situated risks such as deception. Based on a systematic review of 176 studies, we propose the first four-dimensional analytical framework for socially intelligent multimodal systems, critically exposing technical path biases. We advocate for next-generation models that are socially aware, interactively capable, and ethically aligned. Accordingly, we introduce a social competency evaluation suite and a human-centered assessment agenda, advancing multimodal AI from perceptual recognition toward genuine social understanding.
This study investigates the capability of large language models (LLMs) to identify 20 clinically relevant social signals—such as physician dominance and patient affiliation—solely from transcribed clinical dialogue. Method: We introduce the first LLM-based framework capable of concurrently tracking all 20 manually annotated social behaviors, integrating task-specific prompt engineering, comparative evaluation across multiple architectures (GPT and Llama series), and clinical-context-aware prompt optimization, rigorously evaluated on highly imbalanced real-world clinical data. Contribution/Results: Results demonstrate that LLMs can reliably infer nonverbal social behaviors from text alone; prompt design and domain adaptation to healthcare significantly improve classification accuracy; and systematic analysis reveals intrinsic contextual sensitivity patterns in model behavior. This work establishes a reproducible methodological foundation and empirical evidence for automated, fine-grained analysis of clinical communication.
Existing approaches struggle to effectively detect diverse forms of social interaction—including face-to-face, virtual, and hybrid modalities—in naturalistic everyday settings, often being constrained to controlled environments or relying on strong assumptions. This work proposes a smartwatch-based, on-device multimodal system that integrates foreground speech recognition with context-aware modeling to enable, for the first time, real-time dynamic detection of varied social interactions in unconstrained, real-world conditions. Leveraging a speech detector and a multimodal fusion model trained on public datasets, the system identifies 1,691 interaction episodes from over 900 hours of in-the-wild wearable data, with 77.28% validated by users. Evaluated across 33,698 fifteen-second windows, it achieves a balanced accuracy of 90.36% and a sensitivity of 91.17%, overcoming conventional limitations related to fixed temporal windows and multi-speaker scenarios.
Existing datasets struggle to support coupled analysis of affect across individual, interpersonal, and group levels in collaborative settings, and often suffer from fragmented, misaligned multimodal signals. To address this gap, this work introduces a high-ecological-validity multimodal dataset comprising synchronized physiological, eye-tracking, audio, continuous self-reported affect, personality traits, and task performance data from 10 four-person groups (40 participants total) engaged in four distinct collaborative tasks. All signals are temporally aligned and organized following a BIDS-inspired structure with Croissant metadata specifications. The dataset achieves 91% and 98% coverage for physiological and eye-tracking signals, respectively, and includes validation via affect manipulation checks. For the first time, it integrates a three-level analytical framework and provides leave-one-group-out cross-validation baselines alongside 15 reproducible benchmark tasks, offering a standardized, high-coverage resource for group affect research.