Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of AI scientists predominantly rely on synthetic tasks or retrospective objectives, which inadequately capture their capabilities in novelty generation, reasoning, and hypothesis formulation. This work proposes using rapidly evolving real-world adversarial domains—such as Formula 1 car design and Magic: The Gathering deck construction—as dynamic benchmark environments to assess core scientific competencies of AI systems. For the first time, such domains are leveraged to evaluate critical gaps including idea filtering, prioritization, and coherent novelty. Empirical results demonstrate that GPT-5.2 replicates 10 out of 40 documented innovations in F1 design, while Gemini 3 Flash successfully generates 5 of the 7 novel cards in a tournament-winning Magic deck, with card selection preferences strongly aligned with professional metagame trends (Spearman ρ = 0.74), thereby validating the benchmark’s efficacy and real-world alignment.
📝 Abstract
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $ρ= 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.
Problem

Research questions and friction points this paper is trying to address.

AI scientist
benchmarking
novelty
scientific reasoning
hypothesis formulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial real-world domains
AI scientist benchmarking
novelty evaluation
hypothesis formulation
expert alignment