LabFactory: Building and Evaluating Executable AI Labs

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant challenge of automatically constructing executable computational systems from requirement descriptions in scientific tasks. To this end, it proposes AI Builder, a framework that leverages an agent-based architecture, metered workspaces, interface encapsulation, and retrieval-augmented techniques to autonomously transform scientific briefs into executable AI laboratories comprising models, tools, and controllers. The resulting systems are verified in isolation on independent hosts. Crucially, this paradigm evaluates delivered systems rather than procedural descriptions, thereby enabling true end-to-end automated construction. Across 28 construction instances spanning seven task categories, the proposed approach surpasses reference baselines on all 33 sub-tests, demonstrating its effectiveness in automating the generation of functional scientific computing environments.
📝 Abstract
Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how they are used at inference. We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface. The builder develops and packages the lab in a metered workspace; a separate host then executes the delivered artifact on held-out inputs, with reference labels kept outside the solver's input interface, and scores its outputs under the task's protocol. This makes the delivered system, rather than the builder's account of its progress, the object of evaluation. We document 28 selected constructions across seven scientific task categories---from molecular and genomic prediction to physiological signals, clinical decision support, and biomedical text---whose delivered labs exceeded their configured reference values on all 33 subtests under host-side execution. Ten contain predictive models fitted during construction; the others assemble retrieval systems, executable analysis environments, and tool-driven workflows around a fixed platform LLM. Together they show that an AI agent can carry a scientific brief all the way to a working lab that can still be invoked, inspected, and checked after construction ends.
Problem

Research questions and friction points this paper is trying to address.

executable AI labs
scientific task automation
AI agent evaluation
task-specific solver
Innovation

Methods, ideas, or system contributions that make the work stand out.

Executable AI Labs
LabFactory
AI Builder Agent
Host-side Evaluation
Scientific Task Automation
🔎 Similar Papers
No similar papers found.