STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear capacity of vision-language models (VLMs) to comprehend complex spatiotemporal relationships and human intentions in social robot navigation. To this end, it proposes SocialNav-SUB, the first unified evaluation benchmark tailored for this scenario, which systematically assesses the spatial, spatiotemporal, and social reasoning capabilities of VLMs through visual question answering (VQA) tasks against human consensus and rule-based baselines. The findings reveal that while state-of-the-art VLMs approximate human performance, they critically underperform simple rule-based baselines in social reasoning. This work thereby identifies key limitations and delineates clear directions for future improvement, with all data and code released publicly.
📝 Abstract
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/socialnav-sub.
Problem

Research questions and friction points this paper is trying to address.

Social Robot Navigation
Vision-Language Models
Scene Understanding
Spatiotemporal Reasoning
Visual Question Answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Social Robot Navigation
Vision-Language Models
Visual Question Answering
Scene Understanding Benchmark
Spatiotemporal Reasoning
🔎 Similar Papers
No similar papers found.
Nathan Tsoi
Nathan Tsoi
Postdoctoral Researcher, The University of Texas at Austin
Robot LearningSystemsHuman-Robot Interaction
M
Michael J. Munje
Department of Computer Science, The University of Texas at Austin
T
Tejas Oberoi
Department of Computer Science, The University of Texas at Austin
R
Rishab Maheshwari
Department of Computer Science, The University of Texas at Austin
P
Pengen Zheng
Department of Computer Science, The University of Texas at Austin
T
Tanush Chauhan
Department of Computer Science, The University of Texas at Austin
Peter Stone
Peter Stone
Professor of Computer Science, The University of Texas at Austin
Artificial IntelligenceMachine LearningReinforcement LearningMultiagent SystemsRobotics
Joydeep Biswas
Joydeep Biswas
Associate Professor, Computer Science Department, The University of Texas at Austin
RoboticsArtificial IntelligenceMulti Robot SystemsLocalizationMapping