Score
Designs, builds, or evaluates multimodal (vision–language) models and systems that perceive and represent social groups and interactions from combined visual and textual inputs. Uses those systems to infer group membership, spatial arrangements and distances, individual roles and relations, formations, and short-term interaction dynamics, and to analyze how model representations encode social structure.
This paper identifies four critical limitations in current LLM-driven multimodal human behavior understanding systems: (1) overreliance on the “modality-to-text” paradigm, neglecting fine-grained audiovisual social cues; (2) absence of adaptive interactive reasoning capabilities; (3) evaluation confined to static benchmarks, lacking social context and human-centered perspectives; and (4) ethical discourse focused predominantly on legal risks while overlooking socially situated risks such as deception. Based on a systematic review of 176 studies, we propose the first four-dimensional analytical framework for socially intelligent multimodal systems, critically exposing technical path biases. We advocate for next-generation models that are socially aware, interactively capable, and ethically aligned. Accordingly, we introduce a social competency evaluation suite and a human-centered assessment agenda, advancing multimodal AI from perceptual recognition toward genuine social understanding.
This study addresses the limitation of social robots in open-domain dialogue—overreliance on unimodal language models and insufficient visual perception and cross-modal understanding. We propose a multimodal social dialogue framework integrating large language models (LLMs) with vision-language models (VLMs), enabling real-time, synergistic comprehension and generation of visual and linguistic information through environment perception, cross-modal alignment, and dynamic context modeling. Our key contribution is the first systematic investigation of adaptive VLM deployment in general social interaction scenarios, explicitly identifying core technical requirements and challenges for multimodal social dialogue. Experimental results demonstrate significant improvements in dialogue naturalness, situational consistency, and user engagement. The framework establishes a scalable theoretical foundation and practical paradigm for embodied social intelligence.
This work addresses the limited capability of multimodal large language models (MLLMs) in understanding human social interactions and the absence of systematic evaluation frameworks. To bridge this gap, the authors propose Social Caption, a novel benchmark grounded in interaction theory, which introduces the first three-dimensional evaluation framework encompassing social reasoning, holistic social analysis, and targeted social analysis. Through structured tasks and tailored metrics, the framework rigorously assesses models’ social comprehension abilities. Experimental results demonstrate that both model scale and architecture significantly influence social cognition performance, underscoring the framework’s effectiveness and innovation in advancing automated evaluation of social understanding in artificial intelligence systems.
This study addresses the challenge of integrating and analyzing heterogeneous multimodal data—such as video, audio, handwritten notes, and eye-tracking streams—in collaborative design, which often renders collaboration mechanisms and decision-making processes opaque. We propose a modular, extensible multimodal analysis framework that integrates AI-driven artifact auto-extraction, multi-stream temporal alignment graphs, thematic card summarization, and drill-down interactive analysis, implemented in an interactive visual system named reCAPit. Our key contribution lies in enabling semantic-level cross-modal fusion and transparent, interpretable explanations, thereby significantly enhancing process traceability and comprehensibility. Evaluation across six interdisciplinary workshops—including urban planning and ensemble music rehearsal—demonstrates that the framework effectively uncovers collaborative dynamics, supports decision provenance, and improves communication efficiency between researchers and practitioners.
How do language models—lacking explicit visual pretraining—achieve image understanding? Method: We systematically analyze 16 multimodal large language models (MLLMs) spanning four architectural families and four parameter scales. Introducing the concept of “vision-preferring attention heads,” we identify such heads via attention behavior analysis, statistical modeling of attention weights, and cross-scale ablation experiments, empirically validating their strong, consistent focus on visual tokens. Contribution: We are the first to discover and formally define this generalizable, modular visual-perception substructure within LLMs. Our work reveals the pivotal role of attention mechanisms in cross-modal adaptation, demonstrating how vision-preferring heads mediate text–vision alignment. This provides an interpretable, spatially localizable mechanism underlying joint text–vision representation learning, thereby advancing research toward transparent, controllable, and analyzable multimodal foundation models.
This study investigates whether large multimodal models possess human-like theory of mind (ToM) capabilities—specifically, spatiotemporal reasoning about beliefs, intentions, and emotions in dynamic video scenes. To this end, we propose the first end-to-end video-to-text ToM reasoning framework, introducing a novel keyframe retrieval mechanism to explicitly expose the model’s internal reasoning trajectory. We further construct a video-centric ToM benchmark and a probe-based evaluation methodology. Experimental results demonstrate that multimodal large language models exhibit emergent video-based ToM capabilities: they substantially outperform text-only baselines on social-emotional reasoning (+23.6% accuracy) and yield highly interpretable, stepwise reasoning traces grounded in visual evidence. Our work establishes a new paradigm for multimodal cognitive modeling and provides empirical foundations for developing trustworthy, socially aware AI systems.
Short-form videos pose significant challenges for standardized modeling of user engagement due to their multimodal content and platform-specific algorithms. This study addresses this gap by computationally operationalizing classical interpretive theories from narratology, rhetoric, communication studies, and semiotics at scale. Leveraging a multimodal large language model, we automatically annotated 77 theory-driven structural variables across approximately 10,000 TikTok videos from Estonian brands and institutions, supplemented by human validation to assess reliability. Controlling for account size and video age, our model yielded a stable, albeit modest, improvement in predicting user engagement. Results indicate that variables related to perception and communication were reliably annotated, whereas deeper semiotic and archetypal structures proved more challenging to capture. This work establishes a systematic computational framework for analyzing the cultural structures embedded in short-form video content.
This work addresses the limitation of existing social intelligence benchmarks, which predominantly focus on textual modalities while neglecting the critical role of visual cues in social interaction. To bridge this gap, the authors introduce SocialBench, the first multimodal social simulation benchmark comprising 240 scenarios, 585 characters, and 2,340 tasks. SocialBench systematically evaluates agents’ visual social competencies across four hierarchical role-based task categories—facial expression, personality manifestation, interaction regulation, and outcome achievement—leveraging image-text aligned evidence and structured character profiles. Evaluations of multimodal large language models (MLLMs) under both verbalized-vision and direct-vision paradigms reveal that while current models nearly saturate performance in character-specific expression and conflict handling, they exhibit significant deficiencies in interaction regulation and outcome achievement when these tasks rely on visual cues, thereby uncovering substantial challenges in high-level visual social reasoning.
研究探讨语言在多模态模型中的位置,通过分析语言对人类感知和认知的影响,提出语言应作为模型边界和共享代码本存在,而非内部表示。
This study addresses the lack of systematic investigation into the design space of large language model (LLM)-based social simulations, which hinders the assessment of simulation fidelity. It reveals for the first time that this design space exhibits a nontrivial geometric structure. Through systematic analysis of key design choices—including base LLM type and agent connectivity patterns—and their interaction effects, the work identifies the base LLM as the dominant factor shaping simulation outcomes. While some parameters exert additive effects, others engage in complex interactions. Integrating LLM-based agent modeling, social network topology, and survey-validated opinion alignment among agents, this research provides a reproducible design framework for constructing realistic and trustworthy silicon-based societies.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.