🤖 AI Summary
Current embodied intelligence evaluation relies on costly, non-scalable interactive simulations or real-world deployments. To address this, we propose StaticEmbodiedBench—the first plug-and-play benchmark grounded in static scene representations, covering 42 diverse scenes and eight core capability dimensions, enabling non-interactive, large-scale offline evaluation. We design a unified static evaluation framework that integrates vision-language models (VLMs) and vision-language-action models (VLAs) for joint visual–linguistic–action reasoning, and introduce a lightweight interface for automated multi-dimensional assessment. We evaluate 19 VLMs and 11 VLAs, releasing the first static leaderboard for embodied intelligence. Furthermore, we open-source a 200-sample subset and the complete evaluation toolkit to foster standardized, reproducible, and low-cost evaluation paradigms.
📝 Abstract
Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.