π€ AI Summary
This paper identifies four critical limitations in current LLM-driven multimodal human behavior understanding systems: (1) overreliance on the βmodality-to-textβ paradigm, neglecting fine-grained audiovisual social cues; (2) absence of adaptive interactive reasoning capabilities; (3) evaluation confined to static benchmarks, lacking social context and human-centered perspectives; and (4) ethical discourse focused predominantly on legal risks while overlooking socially situated risks such as deception. Based on a systematic review of 176 studies, we propose the first four-dimensional analytical framework for socially intelligent multimodal systems, critically exposing technical path biases. We advocate for next-generation models that are socially aware, interactively capable, and ethically aligned. Accordingly, we introduce a social competency evaluation suite and a human-centered assessment agenda, advancing multimodal AI from perceptual recognition toward genuine social understanding.
π Abstract
LLM-powered multimodal systems are increasingly used to interpret human social behavior, yet how researchers apply the models' 'social competence' remains poorly understood. This paper presents a systematic literature review of 176 publications across different application domains (e.g., healthcare, education, and entertainment). Using a four-dimensional coding framework (application, technical, evaluative, and ethical), we find (1) frequent use of pattern recognition and information extraction from multimodal sources, but limited support for adaptive, interactive reasoning; (2) a dominant 'modality-to-text' pipeline that privileges language over rich audiovisual cues, striping away nuanced social cues; (3) evaluation practices reliant on static benchmarks, with socially grounded, human-centered assessments rare; and (4) Ethical discussions focused mainly on legal and rights-related risks (e.g., privacy), leaving societal risks (e.g., deception) overlooked--or at best acknowledged but left unaddressed. We outline a research agenda for evaluating socially competent, ethically informed, and interaction-aware multi-modal systems.