IMBench: A Benchmark for Intuitive Robotic Manipulation

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate the integrated capabilities of robots across perception, physical reasoning, and action execution, and fail to quantify human-like intuitive manipulation. This work proposes the first evaluation framework that formalizes intuitive manipulation as a cross-modal integration capability, introducing a benchmark comprising 35 tasks and 14K curated trajectories. The benchmark explicitly incorporates contact-rich manipulation, tool use, and multi-stage dependencies, accompanied by an extensible procedural scene generation toolkit. Experiments reveal that current vision-language models lack executable planning abilities, while state-of-the-art vision-language-action models still fall short in satisfying physical constraints and generalizing effectively—highlighting intuitive manipulation as a critical gap in general-purpose robotic policies.
📝 Abstract
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
Problem

Research questions and friction points this paper is trying to address.

intuitive manipulation
physical reasoning
robotic manipulation
benchmark
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

intuitive manipulation
physical reasoning
action generation
robotic benchmark
vision-language-action models
🔎 Similar Papers
No similar papers found.
Anurag Maurya
Anurag Maurya
Indian Institute of Science, Bengaluru
Robot LearningHuman Robot Interaction
S
Sukhvansh Jain
Manav Robotics, Bengaluru, India
P
Prajwal Avhad
Manav Robotics, Bengaluru, India
G
Gautham Balachandran
Manav Robotics, Bengaluru, India
Z
Ziyi Zhou
Manav Robotics, Bengaluru, India
A
Atharva Kshirsagar
Manav Robotics, Bengaluru, India
S
Satyam Singh
Manav Robotics, Bengaluru, India
B
Bowen Li
Manav Robotics, Bengaluru, India
R
Rishabh Mukund
Manav Robotics, Bengaluru, India
R
Ritul Singh
Manav Robotics, Bengaluru, India
J
Jatin Vira
Manav Robotics, Bengaluru, India
S
Suvonil Chatterjee
Manav Robotics, Bengaluru, India
Devesh K. Jha
Devesh K. Jha
Senior Principal Research Scientist, MERL
Machine LearningRoboticsArtificial IntelligenceMotion Planning