EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deficiency of tool-centric reasoning in multimodal models when processing real-world first-person videos by proposing the first comprehensive suite encompassing both data and benchmarks. Methodologically, we construct a large-scale, densely annotated 100-hour first-person tool-use dataset and integrate 3D information to perform supervised fine-tuning on Qwen3-VL. Concurrently, we design a multi-level diagnostic benchmark spanning perceptual, geometric, procedural, and causal reasoning. Experimental results demonstrate that fine-tuning improves model accuracy from 50.0% to 60.9%, effectively validating the utility of the proposed resource suite while revealing critical shortcomings in tool understanding among existing embodied agents.
📝 Abstract
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning about affordances, hand-tool-object geometry, procedural progress, and causal effects on target objects. Yet despite strong performance on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this form of tool-centric embodied reasoning. Progress in this direction has been limited by the lack of real-world egocentric data and diagnostic benchmarks. To address this gap, we introduce EgoTools, the first comprehensive suite for egocentric tool-use understanding. It consists of two complementary components: EgoTools-Data, a large-scale corpus of 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations, and supplementary 3D information; and EgoTools-Bench, a diagnostic benchmark of 1,000 QA pairs across four tracks that cover tool-use understanding from perception and geometry to procedure and causal reasoning. Experimental results show that current models still struggle to ground tool use in visual evidence: Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding. Beyond evaluation, we validate EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation. Together, these results establish EgoTools as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.
Problem

Research questions and friction points this paper is trying to address.

egocentric video
tool-use reasoning
embodied AI
multimodal models
diagnostic benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

egocentric video
tool-use reasoning
diagnostic benchmark
multimodal understanding
embodied AI
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shulin Tian
Shulin Tian
PhD Student, Nanyang Technological University
Computer Vision
J
Junsu Kim
S-Lab, Nanyang Technological University; KAIST
Shuai Liu
Shuai Liu
School of Electrical and Electronic Engineering, Nanyang Technological University
OptimizationOptimal controlTime Delay SystemMulti-agent SystemSignal Processing
H
Hao Li
S-Lab, Nanyang Technological University
Y
Yujiao Shen
S-Lab, Nanyang Technological University
S
Sihan Li
S-Lab, Nanyang Technological University
Z
Zhe Yang
S-Lab, Nanyang Technological University
Y
Yeongon Kim
S-Lab, Nanyang Technological University
F
Feiyu Li
Peking University
J
Jialin Wu
Fudan University
Y
Yichi Zhang
S-Lab, Nanyang Technological University
W
Wenhui Wang
School of Biological Sciences, Nanyang Technological University
R
Runmao Yao
S-Lab, Nanyang Technological University
Yuhao Dong
Yuhao Dong
Tsinghua University, Nanyang Technological University
Multi-modal LearningComputer Vision
Zhaoxi Chen
Zhaoxi Chen
Ph.D. Student, Nanyang Technological University
Neural renderingGenerative models
Fangzhou Hong
Fangzhou Hong
Nanyang Technological University
3D Computer Vision
Antonino Furnari
Antonino Furnari
Assistant Professor at the University of Catania
Computer Vision
Jingkang Yang
Jingkang Yang
PhD, MMLab@NTU
Visual PerceptionVisual ReasoningMultimodalityOpen World
H
Hongyuan Zhu
A*STAR
Ziwei Liu
Ziwei Liu
Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics