social signal processing

Designs, builds, and evaluates algorithms, software, and datasets that detect, extract, fuse, and model human social signals (e.g., facial expressions, gaze, gesture, prosody, posture, turn‑taking) from multimodal interaction data. Uses those components to infer, track, or synthesize interpersonal states, intentions, engagement, rapport, and other social dynamics for real‑time or offline analysis.

socialsignalprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the scarcity of publicly available datasets that simultaneously capture multimodal behavioral signals and self-reported annotations, which has hindered the modeling of stance (agree/disagree/neutral) and social cues in dyadic conversations. To bridge this gap, the authors present a novel multimodal interaction corpus comprising 45 dyads (90 participants), encompassing both familiar and unfamiliar pairs. The dataset integrates synchronized recordings of 2D/3D facial video, thermal imaging, audio, and multiple physiological signals—including photoplethysmography (PPG), electrodermal activity (EDA), heart rate, blood pressure, and respiration—alongside structured stance labels and subjective emotional reports. Spanning approximately 20 terabytes, this resource is the first to jointly provide fine-grained behavioral data and introspective self-assessments, thereby enabling robust computational modeling of interpersonal social dynamics and filling a critical void in the field.

affect annotationconversational stancedyadic interaction

Existing research often examines textual, visual, or audio modalities of short videos in isolation, failing to uncover how their interplay influences user engagement. This work proposes the first reproducible and interpretable multimodal analysis framework that integrates automated feature extraction with Shapley-value-based attribution to systematically investigate how multimodal interactions affect view counts in TikTok content related to social anxiety disorder. The study reveals that facial expressions are more predictive than textual sentiment, that informational content garners greater attention than emotional support, and that multimodal synergies exhibit strong threshold-dependent effects—thereby transcending the limitations of conventional unimodal analyses.

cross-modal interactionmental health discoursemultimodal communication

psifx - Psychological and Social Interactions Feature Extraction Package

Jul 14, 2024
GR
Guillaume Rochette
🏛️ UNIL | University of Surrey

In psychology and social sciences, manual annotation of multimodal behavioral data is costly, time-consuming, and suffers from low inter-annotator consistency. To address this, we present an open-source, modular, task-oriented multimodal behavioral feature extraction toolkit. The toolkit integrates speaker diarization, automatic speech recognition (ASR), machine translation, and vision-based pose estimation (e.g., MediaPipe and OpenPose) into an end-to-end configurable pipeline, enabling automated extraction of speaker identity, transcribed and translated speech, body/gesture/face pose, and gaze trajectories from audiovisual recordings. Its novel extensible architecture significantly lowers the barrier to entry for non-AI-expert researchers, promoting standardization and democratization of behavioral analysis tools. Empirical evaluation demonstrates high accuracy and plug-and-play usability, substantially improving annotation efficiency and consistency. The toolkit has already enabled multiple real-time studies on dynamic human behavior analysis.

Automate and standardize human labor-intensive data annotationDevelop open-source community-driven psychology research softwareEnable large-scale access for non-expert users

Toward Socially-Aware LLMs: A Survey of Multimodal Approaches to Human Behavior Understanding

Oct 27, 2025
ZL
Zihan Liu
🏛️ University of Illinois Urbana-Champaign

This paper identifies four critical limitations in current LLM-driven multimodal human behavior understanding systems: (1) overreliance on the “modality-to-text” paradigm, neglecting fine-grained audiovisual social cues; (2) absence of adaptive interactive reasoning capabilities; (3) evaluation confined to static benchmarks, lacking social context and human-centered perspectives; and (4) ethical discourse focused predominantly on legal risks while overlooking socially situated risks such as deception. Based on a systematic review of 176 studies, we propose the first four-dimensional analytical framework for socially intelligent multimodal systems, critically exposing technical path biases. We advocate for next-generation models that are socially aware, interactively capable, and ethically aligned. Accordingly, we introduce a social competency evaluation suite and a human-centered assessment agenda, advancing multimodal AI from perceptual recognition toward genuine social understanding.

Addressing gaps in ethical considerations and evaluation practices for social AIAnalyzing limitations in adaptive reasoning and nuanced social cue interpretationSurveying multimodal approaches to understand human social behavior using LLMs

Can Language Models Understand Social Behavior in Clinical Conversations?

May 07, 2025
MS
Manas Satish Bedmutha
🏛️ UC San Diego | University of Washington

This study investigates the capability of large language models (LLMs) to identify 20 clinically relevant social signals—such as physician dominance and patient affiliation—solely from transcribed clinical dialogue. Method: We introduce the first LLM-based framework capable of concurrently tracking all 20 manually annotated social behaviors, integrating task-specific prompt engineering, comparative evaluation across multiple architectures (GPT and Llama series), and clinical-context-aware prompt optimization, rigorously evaluated on highly imbalanced real-world clinical data. Contribution/Results: Results demonstrate that LLMs can reliably infer nonverbal social behaviors from text alone; prompt design and domain adaptation to healthcare significantly improve classification accuracy; and systematic analysis reveals intrinsic contextual sensitivity patterns in model behavior. This work establishes a reproducible methodological foundation and empirical evidence for automated, fine-grained analysis of clinical communication.

Assessing LLMs' ability to understand social behaviors in clinical dialoguesAutomating extraction of 20 social signals from patient-provider conversationsEvaluating LLM performance across architectures for healthcare communication analysis

Latest Papers

What's happening recently
View more

Existing approaches struggle to effectively detect diverse forms of social interaction—including face-to-face, virtual, and hybrid modalities—in naturalistic everyday settings, often being constrained to controlled environments or relying on strong assumptions. This work proposes a smartwatch-based, on-device multimodal system that integrates foreground speech recognition with context-aware modeling to enable, for the first time, real-time dynamic detection of varied social interactions in unconstrained, real-world conditions. Leveraging a speech detector and a multimodal fusion model trained on public datasets, the system identifies 1,691 interaction episodes from over 900 hours of in-the-wild wearable data, with 77.28% validated by users. Evaluated across 33,698 fifteen-second windows, it achieves a balanced accuracy of 90.36% and a sensitivity of 91.17%, overcoming conventional limitations related to fixed temporal windows and multi-speaker scenarios.

naturalistic settingson-device detectionsmartwatch

Existing datasets struggle to support coupled analysis of affect across individual, interpersonal, and group levels in collaborative settings, and often suffer from fragmented, misaligned multimodal signals. To address this gap, this work introduces a high-ecological-validity multimodal dataset comprising synchronized physiological, eye-tracking, audio, continuous self-reported affect, personality traits, and task performance data from 10 four-person groups (40 participants total) engaged in four distinct collaborative tasks. All signals are temporally aligned and organized following a BIDS-inspired structure with Croissant metadata specifications. The dataset achieves 91% and 98% coverage for physiological and eye-tracking signals, respectively, and includes validation via affect manipulation checks. For the first time, it integrates a three-level analytical framework and provides leave-one-group-out cross-validation baselines alongside 15 reproducible benchmark tasks, offering a standardized, high-coverage resource for group affect research.

affective computingcollaborative tasksgroup interaction

Hot Scholars

RZ

Renwen Zhang

Assistant Professor, Nanyang Technological University
HCIMental HealthSocial SupportHealth Communication
KS

Koustuv Saha

University of Illinois Urbana-Champaign
Computational Social ScienceSocial ComputingHuman-Centered Machine LearningWellbeing
AX

Amy X. Zhang

Associate Professor, Computer Science & Engineering, University of Washington
social computingHCI
JB

Jonathan Bragg

Allen Institute for AI (AI2)
Artificial IntelligenceHuman-Computer InteractionCrowdsourcing
SL

Shiyang Lai

University of Chicago
computational social sciencecollective intelligenceartificial intelligencecomplex system