RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent challenge of correlating visual observations with interaction histories and task progress in robotic manipulation. To this end, it constructs a streaming task understanding benchmark grounded in real-world robot and human demonstrations, encompassing seven task categories such as recognition, alignment, and temporal localization. Moving beyond aggregate scoring, the work introduces a diagnostic evaluation perspective to conduct zero-shot assessments and input ablation studies across eighteen vision-language models. The findings reveal that strong visual matching capabilities do not inherently translate to temporal ordering proficiency, quantifying the performance gap between them. Furthermore, the analysis demonstrates that next-action prediction relies predominantly on priors rather than real-time visual evidence. Collectively, this research establishes a novel paradigm for evaluating specific model capabilities in embodied AI contexts.
📝 Abstract
Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 vision-language models reveals substantial differences across tasks. GPT-6-Astra achieves 98.3% accuracy on Frame Matching but 68.3% on Frame Ordering, while RynnBrain1.1-122B-A10B exhibits a larger gap, reaching 95.4% and 32.9%, respectively. Input ablations on matched questions with five open-weight models further reveal distinct dependencies on visual evidence: removing visual observations reduces Current Action Recognition accuracy by 22.1 percentage points, whereas Next Action Prediction decreases by only 0.7 points. These findings show that strong visual matching does not consistently coincide with strong temporal ordering, and suggest that next-action prediction can be supported by task and action priors even when visual evidence is unavailable. RoboChrono provides a diagnostic setting for examining these differences, highlighting the need for capability-specific evaluation beyond aggregate scores when assessing task understanding in robot manipulation.
Problem

Research questions and friction points this paper is trying to address.

streaming task understanding
robot manipulation
vision-language models
temporal grounding
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming task understanding
robot benchmark
vision-language models
temporal grounding
input ablation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Yuzhou Wu
Yuzhou Wu
Central South University, Changsha University Of Science And Technology
ICD CODING
Longteng Fan
Longteng Fan
Shanghai Jiao Tong University
computer vision
Z
Zimeng Li
Huazhong Agricultural University
Y
Yu Wanchan
Huazhong Agricultural University
T
Ting Zhang
Shenzhen University
Yiyang Ma
Yiyang Ma
DeepSeek-AI
Generative ModelsLarge Language Models
S
Shihao Li
General Intelligence Machine
W
Wei Ying
South China Agricultural University
Jianbin Qin
Jianbin Qin
Shenzhen University
DatabaseSimilarity SearchMachine Learning
J
Jiajian Jing
Huazhong Agricultural University
F
Fangwen Chen
Huazhong Agricultural University
Y
Yifan Wu
General Intelligence Machine
Zichen Zhang
Zichen Zhang
Virginia Tech
Mechanical Engineering
R
Ruiqi Yang
General Intelligence Machine
W
Weibin Kong
General Intelligence Machine
Y
Yihang Xu
General Intelligence Machine
Haoran Liu
Haoran Liu
Ph.D. Student, Department of Computer Science & Engineering, Texas A&M University
LLMsGraph/Geometric LearningAI for ScienceGenerative Models
Z
Zonghang He
Shanghai Jiao Tong University
Xuyang Liu
Xuyang Liu
Sichuan University
Vision-language ModelsModel CompressionToken CompressionTransfer Learning
Y
YiFan Xiong
Beijing Jiaotong University
Siteng Huang
Siteng Huang
Alibaba DAMO Academy | ZJU | Westlake University
Vision-language ModelsGenerative ModelsEmbodied AI
T
Tao Xu
General Intelligence Machine
Zhuo Xu
Zhuo Xu
Wuhan University
Multi-sensor fusion positioningvisual SLAM
L
Long Chen
General Intelligence Machine
R
Ruoxiang Li
Shenzhen University