agent model discovery

Design and build methods that infer and validate models of agents as stochastic operators mapping task inputs and metrics to outputs, including frameworks that search model space and select models from experimental data. Analyze variability across runs and identify factors or latent variables that affect agent outputs, producing interpretable model representations and uncertainty estimates.

agentmodeldiscovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing benchmarks struggle to capture the stochasticity and adaptability of large language model agents in autonomous model discovery. This work proposes an experimental evaluation paradigm that treats agents as stochastic model discovery operators, systematically examining—through controlled experiments—how task design, objectives, data, and reasoning effort jointly influence output quality, cost, latency, and process complexity. By integrating regression modeling, statistical inference, and utility-aligned decomposition techniques, we evaluate coding agents such as Codex and Claude Code on the WordCraft platform. Our analysis reveals that reasoning effort exerts a dominant and interpretable influence on both performance and cost, thereby demonstrating the framework’s effectiveness and analytical insight.

Agentic AIAutonomous Model DiscoveryExperimental Design

Inferring Capabilities from Task Performance with Bayesian Triangulation

Sep 21, 2023
JB
John Burden
🏛️ University of Cambridge | The Alan Turing Institute | Universitat Politècnica de València

Addressing the challenge of reliably inferring AI systems’ cognitive capabilities from heterogeneous, few-shot task performance, this paper proposes a Bayesian triangulation framework for cognitive profiling. The method introduces a “measurement layout” generative model (implemented in PyMC) that jointly models task-instance features, latent capability dimensions, and system responses—thereby overcoming traditional psychometric reliance on large-scale, homogeneous datasets. Its key innovation lies in the first integration of Bayesian latent-variable modeling with multi-task cross-validation, enabling individualized, architecture-agnostic cognitive capability inversion. Evaluated on the AnimalAI Olympics benchmark (68 competing agents) and the O-PIAAGETS benchmark (30 synthetic agents), the framework successfully reconstructs fine-grained cognitive profiles, significantly enhancing discriminability and interpretability of inferred capabilities. Results empirically validate the feasibility and effectiveness of capability-oriented evaluation as a principled alternative to conventional behavioral benchmarks.

Enable capability inference from non-populational data using Bayesian methodsInfer cognitive profiles from diverse experimental task performance dataModel task-feature and capability interactions affecting system performance

High-dimensional stochastic agent-based models (ABMs) are notoriously difficult to analyze systematically due to the curse of dimensionality and inherent stochasticity. This work proposes a multi-stage automated exploration framework that first employs model-driven experimental design to identify key variables and partition the parameter space, then leverages machine learning surrogate models to efficiently capture residual nonlinear interaction effects. The approach operates without human intervention, automatically detecting unstable regions within the simulator and enabling robust sensitivity analysis and policy testing. Applied to a predator–prey case study, the framework successfully isolates dominant variables and highly sensitive nonlinear regimes, substantially enhancing the efficiency and reliability of ABM exploration.

Agent-Based Modelscurse of dimensionalitynonlinear interactions

This study addresses the lack of systematic evaluation of large language models (LLMs) on stochastic modeling tasks in operations research—spanning theoretical problems (e.g., graduate coursework and Ph.D. qualifying exams) and practical simulation-optimization challenges (drawn from the SimOpt open-source library). Method: We design a multi-tiered benchmark grounded in probability theory, statistics, and stochastic processes to rigorously assess LLMs’ capabilities in uncertainty modeling, analysis, and optimization. Contribution/Results: Experimental results demonstrate that state-of-the-art LLMs achieve near-expert human performance on stochastic modeling tasks, particularly excelling in problem comprehension, model formulation, and solution reasoning. This work establishes the first comprehensive, reproducible evaluation framework for LLMs on operations research problems involving uncertainty, thereby bridging a critical gap in AI assessment and providing empirical foundations for AI-augmented decision modeling.

Assess LLMs' decision-making under uncertainty in OREvaluate LLMs' ability to solve stochastic OR problemsTest LLMs on graduate-level and exam OR problems

Existing evaluation methods fail to quantify AI agents’ capabilities in scientific theory generation, experimental design, and iterative model optimization. To address this, we introduce a generative probabilistic environment benchmark spanning ten scientific domains and propose the first quantifiable, interactive, explanation-driven dual-task evaluation paradigm: (1) experimental design—assessed via Expected Information Gain (EIG) to measure data acquisition quality; and (2) model discovery—evaluated through cross-agent predictive reliability to gauge theoretical explanatory power. Our methodology integrates generative probabilistic modeling, EIG computation, LLM-based scientific agent architectures, and a principled interpretability assessment protocol. Empirical results reveal that state-of-the-art LLMs (e.g., GPT-4o) exhibit significant deficiencies in both core tasks; moreover, augmenting them with explicit statistical modules fails to consistently improve performance—highlighting fundamental limitations of current foundation models in scientific discovery.

Agent Capability AssessmentExperimental Design TestingScientific Theory Evaluation

Latest Papers

What's happening recently
View more

This study addresses the opacity of LLM agent behaviors and the limitations of existing evaluations in elucidating decision logic by pioneering the application of automata learning to strategy discovery. Integrating trajectory abstraction, the proposed method recovers interpretable finite-state models from interaction data, transforming raw trajectories into explicit symbolic representations that facilitate explanation generation and knowledge transfer. Experiments successfully extracted high-level attack strategies from penetration testing agents, effectively supporting systematic interpretability analysis and knowledge distillation for smaller models. Ultimately, this work establishes a novel paradigm for understanding and reusing the intrinsic strategies of intelligent agents through formal symbolic modeling.

agent trajectory analysisbehavioral interpretabilityexplainability

This study addresses the challenge of jointly modeling calibration and control parameters in computer model calibration, where the distribution of calibration parameters is unknown while that of control parameters is known. To tackle this issue, the authors propose a nonparametric Bayesian calibration method based on measure decomposition. The approach preserves the known marginal distribution of the control parameters while employing stochastic process modeling and Bayesian inference to construct a posterior distribution over the input space that aligns with field observations. Notably, this work is the first within a nonparametric calibration framework to explicitly maintain the prior distributional properties of the control parameters, thereby substantially enhancing the physical consistency and scientific credibility of the calibration results.

Calibration ParametersControl ParametersDistribution Preservation

This study addresses the oversight of configuration variability in existing agent evaluations, which leads to significant result variance. Treating agents as configurable systems, this work proposes a systematic evaluation framework that analyzes the effects and interactions of five configuration variables—including information and time—on performance using scientific task benchmarks, alongside a trajectory taxonomy. The findings reveal that specialized validation tools are more effective than prompt engineering at altering agent behavior, and that 54% of performance variance stems from repeated runs. By open-sourcing the evaluation benchmark and a dataset of 18,000 trajectories, this research establishes a new paradigm for the robust evaluation of AI agents.

Agent BehaviorAgentic EvaluationConfigurable Systems

Hot Scholars

GS

Guillaume Sartoretti

Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed Control
HF

Hongwei Feng

Fudan University
knowledge management,AI,big data
GF

Guoliang Fan

Professor of Electrical Engineering at Oklahoma State University
image processingcomputer visionmachine learningmultimedia
YC

Yuhong Cao

National University of Singapore
Robot learningPath Planing