vision-language social reasoning

Designs, builds, or evaluates multimodal (vision–language) models and systems that perceive and represent social groups and interactions from combined visual and textual inputs. Uses those systems to infer group membership, spatial arrangements and distances, individual roles and relations, formations, and short-term interaction dynamics, and to analyze how model representations encode social structure.

vision-languagesocialreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Jul 25, 2025
RJ
Ruben Janssens
🏛️ IDLab-AIRO | Ghent University–imec

This study addresses the limitation of social robots in open-domain dialogue—overreliance on unimodal language models and insufficient visual perception and cross-modal understanding. We propose a multimodal social dialogue framework integrating large language models (LLMs) with vision-language models (VLMs), enabling real-time, synergistic comprehension and generation of visual and linguistic information through environment perception, cross-modal alignment, and dynamic context modeling. Our key contribution is the first systematic investigation of adaptive VLM deployment in general social interaction scenarios, explicitly identifying core technical requirements and challenges for multimodal social dialogue. Experimental results demonstrate significant improvements in dialogue naturalness, situational consistency, and user engagement. The framework establishes a scalable theoretical foundation and practical paradigm for embodied social intelligence.

Adapting vision-language models for autonomous social robotsAddressing technical challenges in multimodal social interactionsEnabling robots to use multiple modalities in social conversations

This work addresses the limited capability of multimodal large language models (MLLMs) in understanding human social interactions and the absence of systematic evaluation frameworks. To bridge this gap, the authors propose Social Caption, a novel benchmark grounded in interaction theory, which introduces the first three-dimensional evaluation framework encompassing social reasoning, holistic social analysis, and targeted social analysis. Through structured tasks and tailored metrics, the framework rigorously assesses models’ social comprehension abilities. Experimental results demonstrate that both model scale and architecture significantly influence social cognition performance, underscoring the framework’s effectiveness and innovation in advancing automated evaluation of social understanding in artificial intelligence systems.

Directed Social AnalysisHolistic Social AnalysisMultimodal Models

A Multimodal Framework for Understanding Collaborative Design Processes

Aug 08, 2025
MK
Maurice Koch
🏛️ University of Stuttgart

This study addresses the challenge of integrating and analyzing heterogeneous multimodal data—such as video, audio, handwritten notes, and eye-tracking streams—in collaborative design, which often renders collaboration mechanisms and decision-making processes opaque. We propose a modular, extensible multimodal analysis framework that integrates AI-driven artifact auto-extraction, multi-stream temporal alignment graphs, thematic card summarization, and drill-down interactive analysis, implemented in an interactive visual system named reCAPit. Our key contribution lies in enabling semantic-level cross-modal fusion and transparent, interpretable explanations, thereby significantly enhancing process traceability and comprehensibility. Evaluation across six interdisciplinary workshops—including urban planning and ensemble music rehearsal—demonstrates that the framework effectively uncovers collaborative dynamics, supports decision provenance, and improves communication efficiency between researchers and practitioners.

Addressing data heterogeneity in workshop observations and outcomesAnalyzing collaborative design processes with multimodal data integrationEnhancing workshop findings through AI extraction and visual analysis

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Dec 24, 2024
JB
Jing Bi
🏛️ University of Rochester | Corning Inc.

How do language models—lacking explicit visual pretraining—achieve image understanding? Method: We systematically analyze 16 multimodal large language models (MLLMs) spanning four architectural families and four parameter scales. Introducing the concept of “vision-preferring attention heads,” we identify such heads via attention behavior analysis, statistical modeling of attention weights, and cross-scale ablation experiments, empirically validating their strong, consistent focus on visual tokens. Contribution: We are the first to discover and formally define this generalizable, modular visual-perception substructure within LLMs. Our work reveals the pivotal role of attention mechanisms in cross-modal adaptation, demonstrating how vision-preferring heads mediate text–vision alignment. This provides an interpretable, spatially localizable mechanism underlying joint text–vision representation learning, thereby advancing research toward transparent, controllable, and analyzable multimodal foundation models.

Analyzing correlation between attention mechanisms and visual understanding capabilitiesIdentifying attention heads specialized in processing visual tokens in multimodal modelsInvestigating how language models interpret visual content without visual training

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Jun 19, 2024
ZC
Zhawnen Chen
🏛️ University of Virginia | The Pennsylvania State University | Northeastern University | Stanford University

This study investigates whether large multimodal models possess human-like theory of mind (ToM) capabilities—specifically, spatiotemporal reasoning about beliefs, intentions, and emotions in dynamic video scenes. To this end, we propose the first end-to-end video-to-text ToM reasoning framework, introducing a novel keyframe retrieval mechanism to explicitly expose the model’s internal reasoning trajectory. We further construct a video-centric ToM benchmark and a probe-based evaluation methodology. Experimental results demonstrate that multimodal large language models exhibit emergent video-based ToM capabilities: they substantially outperform text-only baselines on social-emotional reasoning (+23.6% accuracy) and yield highly interpretable, stepwise reasoning traces grounded in visual evidence. Our work establishes a new paradigm for multimodal cognitive modeling and provides empirical foundations for developing trustworthy, socially aware AI systems.

Developing pipeline for theory-of-mind reasoning using video and textEnabling explicit ToM reasoning through key frame retrievalExamining emotional and social reasoning in multimodal video models

Latest Papers

What's happening recently
View more

Short-form videos pose significant challenges for standardized modeling of user engagement due to their multimodal content and platform-specific algorithms. This study addresses this gap by computationally operationalizing classical interpretive theories from narratology, rhetoric, communication studies, and semiotics at scale. Leveraging a multimodal large language model, we automatically annotated 77 theory-driven structural variables across approximately 10,000 TikTok videos from Estonian brands and institutions, supplemented by human validation to assess reliability. Controlling for account size and video age, our model yielded a stable, albeit modest, improvement in predicting user engagement. Results indicate that variables related to perception and communication were reliably annotated, whereas deeper semiotic and archetypal structures proved more challenging to capture. This work establishes a systematic computational framework for analyzing the cultural structures embedded in short-form video content.

computational content analysisengagement modelingmultimodal annotation

This work addresses the limitation of existing social intelligence benchmarks, which predominantly focus on textual modalities while neglecting the critical role of visual cues in social interaction. To bridge this gap, the authors introduce SocialBench, the first multimodal social simulation benchmark comprising 240 scenarios, 585 characters, and 2,340 tasks. SocialBench systematically evaluates agents’ visual social competencies across four hierarchical role-based task categories—facial expression, personality manifestation, interaction regulation, and outcome achievement—leveraging image-text aligned evidence and structured character profiles. Evaluations of multimodal large language models (MLLMs) under both verbalized-vision and direct-vision paradigms reveal that while current models nearly saturate performance in character-specific expression and conflict handling, they exhibit significant deficiencies in interaction regulation and outcome achievement when these tasks rely on visual cues, thereby uncovering substantial challenges in high-level visual social reasoning.

multimodal simulationsocial interactionsocial-agent benchmark

This study addresses the lack of systematic investigation into the design space of large language model (LLM)-based social simulations, which hinders the assessment of simulation fidelity. It reveals for the first time that this design space exhibits a nontrivial geometric structure. Through systematic analysis of key design choices—including base LLM type and agent connectivity patterns—and their interaction effects, the work identifies the base LLM as the dominant factor shaping simulation outcomes. While some parameters exert additive effects, others engage in complex interactions. Integrating LLM-based agent modeling, social network topology, and survey-validated opinion alignment among agents, this research provides a reproducible design framework for constructing realistic and trustworthy silicon-based societies.

agent-based modelingdesign spaceLLM-based social simulations

This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.

binding probleminformation originmultimodal models

Hot Scholars

TD

Tobias Dienlin

University of Zurich
Media PsychologyCommunicationPrivacyWell-Being
MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
M"

Mohammad "Matt" Namvarpour

PhD Student, Drexel University
Artificial IntelligenceConversational User InterfacesHuman-AI Interactions
MH

Mingfei Han

MBZUAI; University of Technology Sydney; Bytedance Seed; MMLab, SIAT
Object RecognitionVideo UnderstandingVision Language ModelsRobotics
TH

Tilo Hartmann

Vrije Universiteit (VU) Amsterdam
Communication ScienceMedia PsychologyVirtual RealityVideo Games