ArtifactArena: Evaluating Models by What They Build in the Physical World

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of evaluation benchmarks for assessing large language models' capabilities in physical-world engineering construction by developing an open-ended embodied intelligence evaluation platform. Methodologically, the platform requires models to autonomously perform hardware-software co-design for robots and execute tasks within a simulated arena. It introduces an Elo-based tournament mechanism combined with multimodal feedback—encompassing text, physics simulation, and gameplay data—to drive iterative optimization strategies, thereby establishing a continuously evolving and non-saturating dynamic benchmark. Furthermore, this work enables real-time competitive ranking of state-of-the-art models, providing the research community with essential infrastructure for persistently measuring the frontiers of open-ended intelligence.
📝 Abstract
To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit \href{https://artifactarena.ai}{https://artifactarena.ai} for more information.
Problem

Research questions and friction points this paper is trying to address.

AI evaluation
physical grounding
open-ended intelligence
hardware-software co-design
robot engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical Grounding
Hardware-Software Co-design
Open-ended Evaluation
Verifier Guided Refinement
Elo Ranking
🔎 Similar Papers
No similar papers found.