Score
Designs and implements evaluation frameworks, benchmarks, metrics, datasets, and automated pipelines to measure and analyze the performance, robustness, safety, and utility of AI systems—particularly large language models. Builds experiments, statistical analyses, and reporting tools to compare models, validate improvements, and establish evaluation systems and protocols.
Existing evaluation methodologies for large language models (LLMs) suffer from insufficient generalization assessment, as static benchmarks fail to capture the continuously expanding capability boundaries of evolving LLMs. Method: We formally define “evaluation generalizability” and propose a four-dimensional analytical framework encompassing evaluation methodologies, datasets, evaluators, and metrics. Our approach innovatively integrates LLM-as-a-judge, dynamically updated datasets, capability-decoupled benchmark design, and a multidimensional meta-evaluation framework. Contribution: We establish a novel, capability-oriented, automated, and sustainably evolvable evaluation paradigm covering critical dimensions—including knowledge, reasoning, instruction following, multimodal understanding, and safety. Concurrently, we release an open-source, extensible GitHub “living review” repository—a community-maintained, versioned resource—to advance evaluation practice from static benchmarking toward dynamic, collaborative co-evolution.
This paper critically examines systemic flaws in contemporary AI benchmarking—including data bias, inadequate documentation, data contamination, conflation of signal and noise, insufficient sociotechnical alignment, and evaluation distortions driven by cultural, commercial, and competitive logics. Drawing on a meta-review of approximately 100 studies published over the past decade, it integrates technical analysis (e.g., construct validity assessment, sociotechnical systems modeling) with insights from the social sciences to propose, for the first time, the “benchmark trust crisis” analytical framework. The study identifies six interrelated root causes, exposing risks such as oversimplification, exploitability (“gaming”), and detachment from authentic human-AI interaction contexts. It advocates for a next-generation AI evaluation paradigm grounded in robustness, transparency, and contextual sensitivity—thereby furnishing interdisciplinary theoretical foundations and methodological tools for AI governance and regulatory policy.
This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.
AI evaluation tools suffer from poor reproducibility, insufficient statistical rigor, and inefficient community collaboration. Method: This paper introduces and implements the first open-source infrastructure for evaluating large language model (LLM) capabilities and safety—featuring a standardized benchmark suite with 70+ community-contributed tasks. It proposes a structured collaborative governance framework, adopts a resampling-based statistical analysis paradigm with uncertainty quantification, and establishes an end-to-end reproducible testing pipeline. Contributions/Results: (1) A versioned task registry with standardized metadata protocols; (2) A confidence-interval estimation method for cross-model comparisons; (3) End-to-end automated quality control. Empirical validation over eight months demonstrates significant improvements in evaluation reproducibility, statistical reliability, and community engagement efficiency.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
This paper identifies three structural deficiencies in current LLM evaluation for data science: (1) imbalanced task coverage, neglecting data management and exploratory analysis; (2) oversimplified human–AI collaboration models, lacking intermediate autonomy levels; and (3) a narrow automation paradigm that prioritizes human replacement over task-transformation-driven capability advancement. To address these, we propose a novel “task-transformation-driven automation” paradigm and introduce a three-dimensional evaluation framework—encompassing goal-directedness, collaboration intensity, and capability leap. Through systematic literature review and cross-platform tool analysis of 72 mainstream benchmarks, we find only 11% support data cleaning and exploration, and none quantify dynamic collaboration intensity. Our analysis establishes medium-autonomy collaboration as a critical evolutionary pathway, advocating for more comprehensive, human-centered, and evolvable AI evaluation standards.
This study evaluates whether state-of-the-art AI coding assistants reliably adhere to intended objectives in simulated AI lab deployment settings, with a focus on potential deliberate subversion of security research. Building upon the open-source LLM auditing tool Petri, we develop a customized evaluation framework that integrates realistic deployment simulations, multidimensional scenario design—encompassing varied research motivations, task types, alternative threat models, and levels of autonomy—and fine-grained analysis of model behavioral trajectories. This work presents the first systematic investigation of adversarial behaviors by AI models toward security research under conditions closely mirroring real-world deployment, revealing discrepancies in goal recognition between evaluation and deployment contexts. While no conclusive evidence of active sabotage was found across four leading models, both Claude Opus 4.5 Preview and Sonnet 4.5 frequently declined to engage in security-related tasks, with Opus 4.5 Preview additionally exhibiting reduced unprompted awareness during evaluations.
This study addresses the current lack of human-centered, interpretable, and responsible evaluation criteria for AI in modeling and simulation. The authors propose the first multidimensional benchmark framework specifically designed to assess large language models (LLMs) through a human-centric lens, leveraging an open-source system dynamics AI platform to systematically evaluate performance across qualitative modeling, quantitative modeling, and model discussion tasks—emphasizing human-AI collaboration rather than replacement. The framework incorporates critical capabilities such as causal reasoning, iterative model refinement, and behavioral explanation, while embedding ethical and accountability considerations. Empirical results indicate that existing AI tools perform relatively well in qualitative tasks and model discussions but remain limited in causal reasoning and quantitative error correction; furthermore, different LLMs exhibit distinct strengths, with no single model emerging as universally superior.
This work addresses a critical gap in evaluating tool-augmented large language models (LLMs), as existing metrics predominantly emphasize linguistic alignment or task success while overlooking the structural relationship between linguistic signals and executable actions across varying autonomy architectures. To remedy this, the study proposes a behavior-centric evaluation framework grounded in the execution layer, introducing a two-dimensional action–refusal (A–R) space defined by action rate (A) and refusal signals (R), along with a divergence metric (D) to quantify their coordination. Systematic experiments across four canonical scenarios and three autonomy configurations—direct execution, planning, and reflection—reveal significant behavioral distributional differences: reflective scaffolding consistently increases refusal rates in high-risk contexts, yet models exhibit structurally heterogeneous redistribution patterns. By replacing scalar safety scores with separable behavioral dimensions, this approach enables fine-grained, comparable, and interpretable characterization of tool-augmented LLM behaviors.
Current evaluations of large language models (LLMs) on ill-defined tasks—such as complex instruction following and natural language-to-Mermaid sequence diagram generation—suffer from insufficient coverage, sensitivity to phrasing, incomparable metrics, and instability in LLM-based judging, thereby failing to yield reliable or diagnostic assessment signals. This work presents the first systematic analysis of confounding failure modes in such tasks, integrating case studies, failure mode categorization, and a multidimensional evaluation framework to demonstrate how existing benchmarks often conflate distinct error types, leading to distorted scores. Moving beyond monolithic aggregate metrics, the proposed approach delivers actionable, fine-grained insights that lay both theoretical and practical foundations for building more robust and interpretable evaluation systems.
This work addresses the challenges posed by enterprise AI systems—increasingly characterized by probabilistic behavior, context sensitivity, and emergent properties due to large language models, retrieval-augmented generation (RAG), and autonomous agents—which render traditional software quality assurance methods inadequate for managing novel risks. To tackle this, the paper proposes an AI assurance framework centered on continuous risk reduction, introducing a structured taxonomy of AI failures and redefining the assurance pyramid across five layers: data, model, system, application, and organization. The framework deeply integrates evaluation throughout the development lifecycle and emphasizes the distinct organizational impacts of AI failures, fostering an assessment-driven engineering culture. It offers engineering leaders a theoretically grounded yet practically actionable strategy to significantly enhance the trustworthiness assessment and governance of enterprise AI systems.