🤖 AI Summary
This study addresses the challenge of evaluating AI agents in designing physically constructible LEGO brick structures from textual instructions. To this end, it proposes BrickBench, a benchmark that requires agents to retrieve components from a discrete part library, jointly reason over local and global constraints, and generate assembly plans balancing semantic alignment, design aesthetics, and physical feasibility. Technically, the framework integrates multi-scale constraint reasoning with an automated verification environment. The primary contribution is the first comprehensive evaluation system for agent-based brick design encompassing these three dimensions. Experimental results demonstrate that while mainstream agents can satisfy physical and semantic constraints, their overall design proficiency remains significantly inferior to that of human experts.
📝 Abstract
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch