Score
Design and build methods that infer and validate models of agents as stochastic operators mapping task inputs and metrics to outputs, including frameworks that search model space and select models from experimental data. Analyze variability across runs and identify factors or latent variables that affect agent outputs, producing interpretable model representations and uncertainty estimates.
Existing benchmarks struggle to capture the stochasticity and adaptability of large language model agents in autonomous model discovery. This work proposes an experimental evaluation paradigm that treats agents as stochastic model discovery operators, systematically examining—through controlled experiments—how task design, objectives, data, and reasoning effort jointly influence output quality, cost, latency, and process complexity. By integrating regression modeling, statistical inference, and utility-aligned decomposition techniques, we evaluate coding agents such as Codex and Claude Code on the WordCraft platform. Our analysis reveals that reasoning effort exerts a dominant and interpretable influence on both performance and cost, thereby demonstrating the framework’s effectiveness and analytical insight.
Addressing the challenge of reliably inferring AI systems’ cognitive capabilities from heterogeneous, few-shot task performance, this paper proposes a Bayesian triangulation framework for cognitive profiling. The method introduces a “measurement layout” generative model (implemented in PyMC) that jointly models task-instance features, latent capability dimensions, and system responses—thereby overcoming traditional psychometric reliance on large-scale, homogeneous datasets. Its key innovation lies in the first integration of Bayesian latent-variable modeling with multi-task cross-validation, enabling individualized, architecture-agnostic cognitive capability inversion. Evaluated on the AnimalAI Olympics benchmark (68 competing agents) and the O-PIAAGETS benchmark (30 synthetic agents), the framework successfully reconstructs fine-grained cognitive profiles, significantly enhancing discriminability and interpretability of inferred capabilities. Results empirically validate the feasibility and effectiveness of capability-oriented evaluation as a principled alternative to conventional behavioral benchmarks.
High-dimensional stochastic agent-based models (ABMs) are notoriously difficult to analyze systematically due to the curse of dimensionality and inherent stochasticity. This work proposes a multi-stage automated exploration framework that first employs model-driven experimental design to identify key variables and partition the parameter space, then leverages machine learning surrogate models to efficiently capture residual nonlinear interaction effects. The approach operates without human intervention, automatically detecting unstable regions within the simulator and enabling robust sensitivity analysis and policy testing. Applied to a predator–prey case study, the framework successfully isolates dominant variables and highly sensitive nonlinear regimes, substantially enhancing the efficiency and reliability of ABM exploration.
This study addresses the lack of systematic evaluation of large language models (LLMs) on stochastic modeling tasks in operations research—spanning theoretical problems (e.g., graduate coursework and Ph.D. qualifying exams) and practical simulation-optimization challenges (drawn from the SimOpt open-source library). Method: We design a multi-tiered benchmark grounded in probability theory, statistics, and stochastic processes to rigorously assess LLMs’ capabilities in uncertainty modeling, analysis, and optimization. Contribution/Results: Experimental results demonstrate that state-of-the-art LLMs achieve near-expert human performance on stochastic modeling tasks, particularly excelling in problem comprehension, model formulation, and solution reasoning. This work establishes the first comprehensive, reproducible evaluation framework for LLMs on operations research problems involving uncertainty, thereby bridging a critical gap in AI assessment and providing empirical foundations for AI-augmented decision modeling.
Existing evaluation methods fail to quantify AI agents’ capabilities in scientific theory generation, experimental design, and iterative model optimization. To address this, we introduce a generative probabilistic environment benchmark spanning ten scientific domains and propose the first quantifiable, interactive, explanation-driven dual-task evaluation paradigm: (1) experimental design—assessed via Expected Information Gain (EIG) to measure data acquisition quality; and (2) model discovery—evaluated through cross-agent predictive reliability to gauge theoretical explanatory power. Our methodology integrates generative probabilistic modeling, EIG computation, LLM-based scientific agent architectures, and a principled interpretability assessment protocol. Empirical results reveal that state-of-the-art LLMs (e.g., GPT-4o) exhibit significant deficiencies in both core tasks; moreover, augmenting them with explicit statistical modules fails to consistently improve performance—highlighting fundamental limitations of current foundation models in scientific discovery.
This study addresses the opacity of LLM agent behaviors and the limitations of existing evaluations in elucidating decision logic by pioneering the application of automata learning to strategy discovery. Integrating trajectory abstraction, the proposed method recovers interpretable finite-state models from interaction data, transforming raw trajectories into explicit symbolic representations that facilitate explanation generation and knowledge transfer. Experiments successfully extracted high-level attack strategies from penetration testing agents, effectively supporting systematic interpretability analysis and knowledge distillation for smaller models. Ultimately, this work establishes a novel paradigm for understanding and reusing the intrinsic strategies of intelligent agents through formal symbolic modeling.
研究如何从大量开源模型中选择最优模型以设计多智能体系统,通过评估8种模型选择策略,发现单一模型家族内选择表现最佳。
研究提出科学沙箱框架,通过实验、反馈和假设修正循环评估AI代理的科学能力,特别是在生物学模型中优化指标与理解系统规则之间的差异。
This study addresses the challenge of jointly modeling calibration and control parameters in computer model calibration, where the distribution of calibration parameters is unknown while that of control parameters is known. To tackle this issue, the authors propose a nonparametric Bayesian calibration method based on measure decomposition. The approach preserves the known marginal distribution of the control parameters while employing stochastic process modeling and Bayesian inference to construct a posterior distribution over the input space that aligns with field observations. Notably, this work is the first within a nonparametric calibration framework to explicitly maintain the prior distributional properties of the control parameters, thereby substantially enhancing the physical consistency and scientific credibility of the calibration results.
This study addresses the oversight of configuration variability in existing agent evaluations, which leads to significant result variance. Treating agents as configurable systems, this work proposes a systematic evaluation framework that analyzes the effects and interactions of five configuration variables—including information and time—on performance using scientific task benchmarks, alongside a trajectory taxonomy. The findings reveal that specialized validation tools are more effective than prompt engineering at altering agent behavior, and that 54% of performance variance stems from repeated runs. By open-sourcing the evaluation benchmark and a dataset of 18,000 trajectories, this research establishes a new paradigm for the robust evaluation of AI agents.