Toward Socially-Aware LLMs: A Survey of Multimodal Approaches to Human Behavior Understanding

πŸ“… 2025-10-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This paper identifies four critical limitations in current LLM-driven multimodal human behavior understanding systems: (1) overreliance on the β€œmodality-to-text” paradigm, neglecting fine-grained audiovisual social cues; (2) absence of adaptive interactive reasoning capabilities; (3) evaluation confined to static benchmarks, lacking social context and human-centered perspectives; and (4) ethical discourse focused predominantly on legal risks while overlooking socially situated risks such as deception. Based on a systematic review of 176 studies, we propose the first four-dimensional analytical framework for socially intelligent multimodal systems, critically exposing technical path biases. We advocate for next-generation models that are socially aware, interactively capable, and ethically aligned. Accordingly, we introduce a social competency evaluation suite and a human-centered assessment agenda, advancing multimodal AI from perceptual recognition toward genuine social understanding.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Social Cognition And InteractionComputer Vision: Multi-modal Vision

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
πŸ“ Abstract
LLM-powered multimodal systems are increasingly used to interpret human social behavior, yet how researchers apply the models' 'social competence' remains poorly understood. This paper presents a systematic literature review of 176 publications across different application domains (e.g., healthcare, education, and entertainment). Using a four-dimensional coding framework (application, technical, evaluative, and ethical), we find (1) frequent use of pattern recognition and information extraction from multimodal sources, but limited support for adaptive, interactive reasoning; (2) a dominant 'modality-to-text' pipeline that privileges language over rich audiovisual cues, striping away nuanced social cues; (3) evaluation practices reliant on static benchmarks, with socially grounded, human-centered assessments rare; and (4) Ethical discussions focused mainly on legal and rights-related risks (e.g., privacy), leaving societal risks (e.g., deception) overlooked--or at best acknowledged but left unaddressed. We outline a research agenda for evaluating socially competent, ethically informed, and interaction-aware multi-modal systems.
Problem

Research questions and friction points this paper is trying to address.

Surveying multimodal approaches to understand human social behavior using LLMs
Analyzing limitations in adaptive reasoning and nuanced social cue interpretation
Addressing gaps in ethical considerations and evaluation practices for social AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic review of multimodal systems for social behavior
Identifies modality-to-text pipelines limiting social cues
Proposes research agenda for socially-aware ethical systems
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Z
Zihan Liu
University of Illinois Urbana-Champaign, USA
P
Parisa Rabbani
University of Illinois Urbana-Champaign, USA
V
Veda Duddu
University of Illinois Urbana-Champaign, USA
K
Kyle Fan
University of Illinois Urbana-Champaign, USA
M
Madison Lee
University of Illinois Urbana-Champaign, USA
Y
Yun Huang
University of Illinois Urbana-Champaign, USA