Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge in robotic manipulation where current observations are insufficient to determine the state, and existing models struggle to leverage historical interactions for inferring task semantics. To this end, we introduce GiT, a multi-source dataset and evaluation benchmark encompassing scenarios such as biological laboratories. Methodologically, we incorporate counterfactual task pairs and fine-grained subtask annotations to precisely assess the utilization of historical information. The dataset integrates real-robot trajectories, UMI demonstrations, and ManiSkill simulations to facilitate the exploration of Vision-Language-Action (VLA) models. Evaluations reveal that current VLA models remain significantly deficient in manipulation tasks requiring historical context. Ultimately, this work provides critical data support and an evaluation paradigm for advancing history-aware robotic manipulation.
πŸ“ Abstract
Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.
Problem

Research questions and friction points this paper is trying to address.

Temporal Grounding
Robotic Manipulation
Vision-Language-Action Models
History-dependent Inference
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Grounding
Vision-Language-Action Models
Counterfactual Task Pairs
Robotic Manipulation
Multi-Source Benchmark
πŸ”Ž Similar Papers
Y
Yi Wang
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
Y
Yang Yang
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
G
Guangqi Xu
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
S
Sumin Lin
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
N
Ning Kang
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
P
Pengxiang Lu
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
X
Xiaotong Chen
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
Z
Zeyu Xue
H.I.-Robot Lab., School of Automation Science and Engineering, South China University of Technology
P
Ping Deng
Faculty of Mechanical and Aerospace Engineering, The Hong Kong University of Science and Technology, HKSAR
X
Xing Liu
School of Astronautics, Northwestern Polytechnical University, Xi’an, China
Chenguang Yang
Chenguang Yang
Chair Professor in Robotics, Fellow of IEEE, IET, IMechE, AIAA, BCS
Robotics
Zhenyu Lu
Zhenyu Lu
Professor, South China University of Technology
teleoperationrobot controlrobot skill learning