One Run Is Not an Idea: The Implementation Lottery in Automated Research

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the common misuse in automated research of evaluating an idea based on a single experimental run, which is susceptible to “implementation lottery” and can yield unreliable conclusions. To mitigate this, the authors propose an “idea reliability audit” framework that systematically quantifies the impact of implementation variance by conducting multiple independent realizations of the same idea. Leveraging techniques such as candidate card freezing, unbiased fidelity labeling, cross-session re-implementation, and artifact re-execution—alongside novel metrics including intraclass correlation coefficient (ICC) and leave-one-out (LOO) winner reversal rate—the study analyzes 312 tasks and demonstrates that implementation variance significantly exceeds run-to-run variance. Crucially, reliance on a single implementation leads to decision reversals in 25.6%–43.6% of cases, providing the first empirical validation of the necessity and effectiveness of multi-implementation verification.
📝 Abstract
Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.
Problem

Research questions and friction points this paper is trying to address.

implementation lottery
idea reliability
automated research
experimental variance
research evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

implementation lottery
idea reliability
Idea Reliability Audit
winner reversal
automated research