RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决现有SPI研究在动态视频分析和场景覆盖上的不足,本文引入了RobotEQ-Video基准,采用层次化世界状态分类法,推动从静态图像到动态视频的转变。
📝 Abstract
Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.
Problem

Research questions and friction points this paper is trying to address.

Social Proactive Intelligence
dynamic videos
world-state taxonomy
scenario coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Social Proactive Intelligence
video-centric analysis
world-state taxonomy
comprehensive scenario coverage
🔎 Similar Papers
No similar papers found.
X
Xinyi Che
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
Zheng Lian
Zheng Lian
Associate Professor, IEEE/CCF Senior Member, Institute of Automation, Chinese Academy of Sciences
Affective ComputingSentiment AnalysisMachine Learning
K
Kuofei Fang
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
Xuehao Wang
Xuehao Wang
Zhejiang University
Multi-Task LearningSegment Anything ModelPEFTLLM
X
Xinghai Gao
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
J
Junqing Wu
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
C
Chuyu Wu
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
L
Liyi Liu
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
Y
Yanhan Huang
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
K
Keyi Xie
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
H
Haomin Ouyang
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University
J
Jinyang Wu
Tsinghua University
Fan Zhang
Fan Zhang
CSE PhD Student, The Chinese University of Hong Kong (CUHK)
Large Language ModelsAI for ScienceMultimodal Learning
R
Runhao Zeng
Shenzhen MSU-BIT University
X
Xun Yang
University of Science and Technology of China
B
Bin He
State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University