Static and Plugged: Make Embodied Evaluation Simple

📅 2025-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current embodied intelligence evaluation relies on costly, non-scalable interactive simulations or real-world deployments. To address this, we propose StaticEmbodiedBench—the first plug-and-play benchmark grounded in static scene representations, covering 42 diverse scenes and eight core capability dimensions, enabling non-interactive, large-scale offline evaluation. We design a unified static evaluation framework that integrates vision-language models (VLMs) and vision-language-action models (VLAs) for joint visual–linguistic–action reasoning, and introduce a lightweight interface for automated multi-dimensional assessment. We evaluate 19 VLMs and 11 VLAs, releasing the first static leaderboard for embodied intelligence. Furthermore, we open-source a 200-sample subset and the complete evaluation toolkit to foster standardized, reproducible, and low-cost evaluation paradigms.

Technology Category

Intelligent Robots: Embodied AIComputer Vision: Language and VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.
Problem

Research questions and friction points this paper is trying to address.

Addresses costly fragmented embodied intelligence evaluation methods
Introduces plug-and-play benchmark for unified static assessment
Evaluates diverse VLMs VLAs for scalable leaderboard
Innovation

Methods, ideas, or system contributions that make the work stand out.

Plug-and-play benchmark for unified evaluation
Static scene representations for scalable assessment
First unified static leaderboard for embodied intelligence
🔎 Similar Papers
2024-07-09IEEE/ASME transactions on mechatronicsCitations: 94
💼 Related Jobs
No related jobs found.
Jiahao Xiao
Jiahao Xiao
Shanghai Ai Lab | Shanghai Jiaotong University
Embodied AiLow-level Vision
J
Jianbo Zhang
Shanghai AI Lab
B
BoWen Yan
Shanghai AI Lab
S
Shengyu Guo
Shanghai AI Lab
T
Tongrui Ye
Shanghai AI Lab
K
Kaiwei Zhang
Shanghai AI Lab
Z
Zicheng Zhang
Shanghai AI Lab
X
Xiaohong Liu
Shanghai JiaoTong University
Zhengxue Cheng
Zhengxue Cheng
Assistant Researcher, Shanghai Jiao Tong University
Video and Image CodingComputer VisionImage Quality Assessment
L
Lei Fan
Shanghai JiaoTong University
C
Chuyi Li
Shanghai AI Lab
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays