Score
Designs and implements systematic benchmarks, test suites, evaluation protocols, and metrics to measure the validity, reliability, semantic fidelity, and accountability of interpretability methods and explanations; this includes creating sim‑based, deductive, and oracle-style evaluations, success criteria, and interpretability validation procedures. Builds diagnostic models and error‑mode analyses to quantify variance across architectures and training choices, categorize failure types, measure sensitivity to noise and few‑shot or behavioral changes, and validate or refine interpretability metrics and methods.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.
This study addresses the unclear relationships among information sources, generation strategies, and quality evidence in test case generation using large language models (LLMs). Through a systematic literature review of 95 studies, this work constructs a multidimensional taxonomy and a benchmark analysis framework. Specifically, it proposes a four-dimensional classification system that elucidates how execution feedback influences oracle independence. Furthermore, it establishes a unified theoretical framework connecting the generation process with quality assessment. By identifying independent oracle evaluation as a critical yet underexplored dimension, this research formulates a future agenda centered on rigorous, oracle-independent quality measurement for LLM-generated test cases.
本文提出了一种形式化的语义块模型和执行评判基准来独立评估规范质量,通过结构化表示和机器可验证条件解决规范确定性问题。
This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.
This study addresses the scoring distortion in existing tool-agent benchmarks, which assume correct interface behavior while overlooking defects in underlying implementations. We model tool interfaces as executable contracts and propose an auditing framework that integrates static code analysis with dynamic state tracking to systematically verify consistency between tool implementations and interface declarations, thereby tracing the origins of scoring discrepancies. Evaluations across four mainstream benchmarks reveal seven latent defects, demonstrating the predominance of static analysis in defect detection and the limitations of dynamic verification. Furthermore, our findings expose misjudgments in current evaluators, showing that most reported scores fail to reflect actual state changes. This work establishes a new paradigm for enhancing the reliability of agent evaluation.
This study addresses the poor reproducibility, lack of auditability, and reliance on manual narratives in AI safety assessments by proposing a deterministic, auditable framework. The framework standardizes heterogeneous engineering evidence into control identifiers mapped to technical-level risks, generates executable assessment functions by compiling MITRE ATLAS rules, supports repeated evaluations via versioned policy objects, and incorporates formal verification to ensure logical consistency and semantic correctness. Experiments across five open-source projects demonstrate that the framework effectively quantifies risk variations before and after hardening interventions. Results confirm that strengthened controls reduce attack feasibility while precisely revealing residual risks arising from missing core safeguards.
论文针对代理基准测试中的双重测量混淆问题,通过将关键决策转移给模型、使用基于真实值的评分及报告更全面的可靠性指标来解决。
研究审计了八个网络安全基准在不同语言模型上的表现,揭示了评分依赖于评估流程配置的问题,并提出标准化评估流程以提高模型评价可靠性。
Current agent benchmarks rely on manual auditing, which struggles to scale and often fails to identify validity flaws, thereby undermining the credibility of model capability evaluations. This work proposes the first automated AI scanner tailored for agent benchmarking, leveraging large language models and structured scoring rules to detect four categories of validity issues in agent transcripts. The system is calibrated through human annotation validation and cross-benchmark evaluation. Experiments across five prominent benchmarks uncover multiple quality issues that evade manual spot-checking, demonstrating that the proposed method effectively enables systematic auditing of benchmarks. This approach establishes a new paradigm for enhancing the reliability of agent evaluations.