agent experiment design

Designs controlled experiments to evaluate agent and multi-agent systems by specifying experimental factors, randomized runs and replications, task assignments, and the task-metric responses to collect; and builds data-collection protocols and statistical analysis plans (e.g., regression models and related analyses) to estimate effects and compare agent behaviors.

agentexperimentdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of effectively evaluating the impact of AI systems in knowledge work, which is hindered by traditional experimental methods that rely on unstructured textual descriptions lacking comparability, reusability, and auditability. To overcome this limitation, the authors propose the SEED framework, which formalizes human–AI collaborative experimental designs as typed participant–process graphs. This approach enables explicit representation of interaction structures, assessment of design novelty, and generation of feasible configurations under specified constraints. Integrating structured encoding, graph-guided generation, and lightweight validation, SEED significantly enhances process clarity, hypothesis specificity, and regulatory compliance in a medical triage task. The results demonstrate its effectiveness as a traceable, comparable, and generative tool for supporting rigorous experimental design in human–AI collaboration.

AI governanceexperimental designhuman-AI collaboration

PLanet: Formalizing Experimental Design

May 14, 2025
LB
London Bielicke
🏛️ UCLA | MIT | University of Massachusetts Amherst | Amazon Web Services | University of California, Los Angeles

Experimental designs in scientific papers often lack clarity, communicability, and comparability, undermining conclusion reliability and generalizability. To address this, we propose the first composable, formal syntax framework for experimental design—realized as a domain-specific language (DSL)—that supports three-stage modeling: experimental unit definition, trial sequence generation, and mapping. The DSL explicitly encodes implicit design decisions (e.g., Latin square allocation), enabling precise specification and reasoning about experimental structure. This framework fills a critical formalization gap in human-computer interaction and related empirical disciplines. We empirically evaluated it on 12 studies from CHI and UIST, successfully formalizing 11. Our analysis uncovered previously unstated design ambiguities and viable alternatives, thereby enhancing experimental transparency, reproducibility, and cross-study comparability.

Challenges in communicating and comparing design alternativesDifficulties in specifying experimental design plans clearlyLack of formal language for constructing experimental assignment procedures

This study addresses the low automation level and high human dependency in scientific research workflows by proposing an Autonomous Simulation Agent (ASA) framework tailored for long-duration simulation tasks. Methodologically, the ASA integrates prompt engineering, automated code generation, remote high-performance computing (HPC) job scheduling, and multi-stage workflow orchestration, featuring a dynamically self-verifying architecture. A novel local-attention–global-supervision coordination mechanism enables 20 rounds of fully autonomous, human-free iteration. Evaluated on polymer chain conformational sampling, ASA-GPT-4o achieves near 100% task completion rate and sustains stable end-to-end operation across 20 consecutive cycles. The framework significantly enhances research efficiency, operational reliability, and experimental reproducibility, advancing the automation and robustness of computational science workflows.

Efficiency EnhancementLarge Language ModelsScientific Research Automation

Latest Papers

What's happening recently
View more

This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.

evidence-driven replanninglong-horizon planningphysical executability

This study addresses the evaluation and generalization capabilities of large language model (LLM) agents in microscope control tasks by introducing a benchmark framework comprising 53 tasks along with trajectory logs. The authors systematically evaluate 105 configurations—spanning single- to triple-agent topologies, five LLMs, RAG parameters, and operational constraints—based on 1,949 experimental runs and 49,109 RAG retrievals, quantifying differences in latency, token consumption, cost, and failure modes. Their analysis reveals, for the first time, the critical influence of agent architecture on performance. While the benchmark proves effective for certification and cross-configuration comparison, it demonstrates limited reliability in predicting performance on unseen tasks, thereby highlighting the substantial challenge of evaluating generalization in LLM-based agent systems.

agentic microscopybenchmarkinggeneralization

This study addresses a critical gap in understanding how task-oriented Agent Plan artifacts in open-source software guide AI-powered coding tools. For the first time, it systematically identifies and analyzes real-world Agent Plan files from open-source projects by screening 36,710 GitHub repositories and conducting qualitative content analysis focused on Markdown-formatted planning documents. The investigation yields 85 valid Agent Plan files that span key engineering activities—including maintenance, design, and implementation—and explicitly articulate task intent while providing concrete execution steps and validation criteria. These findings reveal the instrumental role such plans play in facilitating human-AI collaborative development and underscore their practical value in structuring and communicating software engineering tasks.

Agent PlansAgentic AI Coding ToolsExecution Guidance

Hot Scholars

YZ

Yi Zhang

Principal Applied Scientist, AWS Agentic AI Labs
Natural Language ProcessingComputational LinguisticsSyntaxParsing
YD

Yali Du

Turing Fellow, Associate professor, King's College London
Multi-Agent Reinforcement LearningHuman-ai coordinationAlignmentCooperative AI
SK

Sergey Kovalchuk

ITMO University
artificial intelligencehuman-AI interactioncomplex systemscomputational science
ZJ

Zhijing Jin

Max Planck Institute
Natural Language ProcessingCausal InferenceMachine LearningArtificial Intelligence
YW

Yu-Wing Tai

Dartmouth College
Computer VisionDeep LearningMulti-modalities Generative AI