Score
Designs and implements tests, benchmarks, probes, and analysis methods to measure and characterize what computational models can do, including task-level performance, generalization, robustness, failure modes, and emergent behaviors. Builds metrics, evaluation protocols, and interpretive analyses that quantify capabilities and limitations across inputs, tasks, and model configurations.
This work addresses the critical gap between the widespread deployment of AI models and the limited understanding of their internal mechanisms, as conventional benchmarks often fail to uncover root causes of failures such as hallucination and shortcut learning. It proposes the first systematic “model science” framework, integrating paradigms from cognitive science, neuroscience, and related disciplines to enable in-depth analysis of individual model instances through four complementary lenses: Verify, Explore, Steer, and Refine. By establishing a shared knowledge repository and collaborative research infrastructure, the framework transcends the limitations of population-level performance evaluation, offering both theoretical foundations and practical pathways to enhance AI interpretability, reliability, and continuous improvement. This paradigm shift moves AI research beyond performance-centric metrics toward a deeper, understanding-driven approach.
Configuration space explosion complicates performance impact modeling, while gray-box approaches rely on structural knowledge (e.g., module execution graphs) to improve model accuracy—yet the mechanisms by which structural features (e.g., number of modules or configuration options) and structural knowledge influence modeling difficulty and optimization potential remain unclear. Method: We formally define “modeling hardness” and “improvement opportunity,” establishing an analytical framework and matrix to quantify the interplay among system structural complexity, structural knowledge level, and modeling benefit. Controlled experiments on synthetic systems integrate module execution graph analysis with gray-box modeling. Contribution/Results: We identify module count and configuration option count as dominant determinants of modeling hardness. Under high hardness, strong structural knowledge significantly increases improvement opportunity. Structural knowledge primarily enhances ranking accuracy, whereas hardness predominantly degrades prediction accuracy. Our findings provide theoretical foundations and strategic guidance for allocating structural knowledge investment according to specific modeling objectives.
Addressing the challenge of reliably inferring AI systems’ cognitive capabilities from heterogeneous, few-shot task performance, this paper proposes a Bayesian triangulation framework for cognitive profiling. The method introduces a “measurement layout” generative model (implemented in PyMC) that jointly models task-instance features, latent capability dimensions, and system responses—thereby overcoming traditional psychometric reliance on large-scale, homogeneous datasets. Its key innovation lies in the first integration of Bayesian latent-variable modeling with multi-task cross-validation, enabling individualized, architecture-agnostic cognitive capability inversion. Evaluated on the AnimalAI Olympics benchmark (68 competing agents) and the O-PIAAGETS benchmark (30 synthetic agents), the framework successfully reconstructs fine-grained cognitive profiles, significantly enhancing discriminability and interpretability of inferred capabilities. Results empirically validate the feasibility and effectiveness of capability-oriented evaluation as a principled alternative to conventional behavioral benchmarks.
This study addresses the lack of capability assessment for large language models (LLMs) in elementary-school visual programming computational thinking evaluation. We introduce the first standardized benchmark specifically designed for this domain, covering multi-level skills—including recognition, selection, and synthesis. To address data scarcity and ensure pedagogical validity, we propose a symbol-rule-based hierarchical synthetic data generation paradigm that explicitly models the developmental progression of computational thinking competencies, which subsequently guides supervised fine-tuning (SFT). Experiments on GPT-4o and Llama3 demonstrate that fine-tuned models significantly outperform baselines on elementary computational thinking assessments, achieving performance comparable to the average human student score. To foster reproducibility and community advancement, we fully open-source all benchmark datasets, symbolic generation rules, and training code—establishing a foundational resource for educational AI evaluation and domain adaptation.
Existing benchmarks struggle to capture the stochasticity and adaptability of large language model agents in autonomous model discovery. This work proposes an experimental evaluation paradigm that treats agents as stochastic model discovery operators, systematically examining—through controlled experiments—how task design, objectives, data, and reasoning effort jointly influence output quality, cost, latency, and process complexity. By integrating regression modeling, statistical inference, and utility-aligned decomposition techniques, we evaluate coding agents such as Codex and Claude Code on the WordCraft platform. Our analysis reveals that reasoning effort exerts a dominant and interpretable influence on both performance and cost, thereby demonstrating the framework’s effectiveness and analytical insight.
This study addresses a critical gap in existing scientific data analysis benchmarks, which fail to differentiate models’ capabilities across distinct scientific reasoning tasks—such as hypothesis exploration, causal inference, and mechanistic explanation. To this end, the authors introduce SDABench, the first multidimensional evaluation benchmark specifically designed to assess scientific analytical competence. It encompasses six dimensions: descriptive, exploratory, inferential, predictive, causal, and mechanistic reasoning, comprising 527 real-world and 6,000 synthetically generated data instances across five scientific domains. Using a five-stage error analysis framework, the benchmark systematically evaluates 15 prominent large language models. Results reveal strong performance on descriptive tasks but substantial deficiencies in complex reasoning involving hypothesis selection, latent variable modeling, and mechanistic inference, indicating that current models remain ill-equipped to support high-level scientific discovery.
This work addresses the limitation of existing benchmarks, which assess only the correctness of agent outputs and fail to capture differences in behavioral strategies. To overcome this, the authors introduce the concept of a “behavioral fingerprint” and present the ProcGrep framework, which leverages emergent lexicon induction to construct compact yet expressive procedural representations of agent trajectories. Behavioral similarity is quantified using Jensen–Shannon divergence, enabling, for the first time, trajectory-based agent identification and cross-task behavioral pattern analysis. Evaluated on SWE-Bench, the method achieves 85.7% accuracy in trajectory attribution and reveals that agents derived from the same origin or training epoch exhibit highly similar behaviors. This approach establishes a new paradigm for agent routing, monitoring, and evolutionary analysis.
This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.