Score
Designs controlled experiments to evaluate agent and multi-agent systems by specifying experimental factors, randomized runs and replications, task assignments, and the task-metric responses to collect; and builds data-collection protocols and statistical analysis plans (e.g., regression models and related analyses) to estimate effects and compare agent behaviors.
This study addresses the challenge of effectively evaluating the impact of AI systems in knowledge work, which is hindered by traditional experimental methods that rely on unstructured textual descriptions lacking comparability, reusability, and auditability. To overcome this limitation, the authors propose the SEED framework, which formalizes human–AI collaborative experimental designs as typed participant–process graphs. This approach enables explicit representation of interaction structures, assessment of design novelty, and generation of feasible configurations under specified constraints. Integrating structured encoding, graph-guided generation, and lightweight validation, SEED significantly enhances process clarity, hypothesis specificity, and regulatory compliance in a medical triage task. The results demonstrate its effectiveness as a traceable, comparable, and generative tool for supporting rigorous experimental design in human–AI collaboration.
Experimental designs in scientific papers often lack clarity, communicability, and comparability, undermining conclusion reliability and generalizability. To address this, we propose the first composable, formal syntax framework for experimental design—realized as a domain-specific language (DSL)—that supports three-stage modeling: experimental unit definition, trial sequence generation, and mapping. The DSL explicitly encodes implicit design decisions (e.g., Latin square allocation), enabling precise specification and reasoning about experimental structure. This framework fills a critical formalization gap in human-computer interaction and related empirical disciplines. We empirically evaluated it on 12 studies from CHI and UIST, successfully formalizing 11. Our analysis uncovered previously unstated design ambiguities and viable alternatives, thereby enhancing experimental transparency, reproducibility, and cross-study comparability.
This study addresses the low automation level and high human dependency in scientific research workflows by proposing an Autonomous Simulation Agent (ASA) framework tailored for long-duration simulation tasks. Methodologically, the ASA integrates prompt engineering, automated code generation, remote high-performance computing (HPC) job scheduling, and multi-stage workflow orchestration, featuring a dynamically self-verifying architecture. A novel local-attention–global-supervision coordination mechanism enables 20 rounds of fully autonomous, human-free iteration. Evaluated on polymer chain conformational sampling, ASA-GPT-4o achieves near 100% task completion rate and sustains stable end-to-end operation across 20 consecutive cycles. The framework significantly enhances research efficiency, operational reliability, and experimental reproducibility, advancing the automation and robustness of computational science workflows.
This study evaluates the reliability and adaptability of large language models in executing scientific tasks within real-world physical environments, with a focus on their ability to generate executable experimental protocols and iteratively refine them based on empirical evidence. Leveraging a robotic chemistry laboratory comprising 45 modular workstations and conducting 4,608 trials, this work extends scientific agent evaluation beyond pure reasoning to encompass physical executability and evidence-driven closed-loop adaptation, introducing a quantifiable framework for assessing deployment readiness. Results reveal that only 3.3% of generated protocols were deemed executable by expert reviewers, with the best-performing system achieving a success rate of 28.1%. Most generated workflows contained no more than 30 steps and generally lacked capabilities for workflow-level replanning or methodological reconfiguration in response to experimental outcomes.
This study addresses the evaluation and generalization capabilities of large language model (LLM) agents in microscope control tasks by introducing a benchmark framework comprising 53 tasks along with trajectory logs. The authors systematically evaluate 105 configurations—spanning single- to triple-agent topologies, five LLMs, RAG parameters, and operational constraints—based on 1,949 experimental runs and 49,109 RAG retrievals, quantifying differences in latency, token consumption, cost, and failure modes. Their analysis reveals, for the first time, the critical influence of agent architecture on performance. While the benchmark proves effective for certification and cross-configuration comparison, it demonstrates limited reliability in predicting performance on unseen tasks, thereby highlighting the substantial challenge of evaluating generalization in LLM-based agent systems.
This study addresses a critical gap in understanding how task-oriented Agent Plan artifacts in open-source software guide AI-powered coding tools. For the first time, it systematically identifies and analyzes real-world Agent Plan files from open-source projects by screening 36,710 GitHub repositories and conducting qualitative content analysis focused on Markdown-formatted planning documents. The investigation yields 85 valid Agent Plan files that span key engineering activities—including maintenance, design, and implementation—and explicitly articulate task intent while providing concrete execution steps and validation criteria. These findings reveal the instrumental role such plans play in facilitating human-AI collaborative development and underscore their practical value in structuring and communicating software engineering tasks.