LabFactory: Building and Evaluating Executable AI Labs
This study addresses the significant challenge of automatically constructing executable computational systems from requirement descriptions in scientific tasks. To this end, it proposes AI Builder, a framework that leverages an agent-based architecture, metered workspaces, interface encapsulation, and retrieval-augmented techniques to autonomously transform scientific briefs into executable AI laboratories comprising models, tools, and controllers. The resulting systems are verified in isolation on independent hosts. Crucially, this paradigm evaluates delivered systems rather than procedural descriptions, thereby enabling true end-to-end automated construction. Across 28 construction instances spanning seven task categories, the proposed approach surpasses reference baselines on all 33 sub-tests, demonstrating its effectiveness in automating the generation of functional scientific computing environments.