Score
Designs and implements evidence-calibration mechanisms and layers that appraise, filter, weight, and align referenced evidence to specified evaluation criteria, including detecting and eliminating unsupported or invalid evidence citations. Builds and analyzes evidence-grounded auditing procedures that produce calibrated confidence scores and precise pass/fail determinations by grounding judgments in validated evidence and reducing unsupported or spurious claims.
This study addresses a critical “evaluation–safety gap” (EvalSafetyGap) in large language models (LLMs), wherein apparent performance gains do not necessarily reflect genuine safety capabilities. Through a systematic literature review, gray literature analysis, and a multidimensional audit of ten models across eight evidence streams, the work proposes the EvalSafetyGap hypothesis and introduces two novel constructs—“instability decomposition” and the “alignment trilemma”—to establish a unified terminology and evidence map supporting dynamic evaluation and auditable alignment. Empirical findings reveal no significant correlation between model capability and adversarial robustness (r = 0.232, p = 0.520). Moreover, safety differences between open- and closed-source models stem primarily from governance transparency rather than behavioral robustness, with results highly sensitive to model categorization and evaluation protocols.
This study addresses the challenge in IT auditing where evidence from heterogeneous organizations is fragmented and compliance with security and regulatory controls must be assessed based on semantic adequacy rather than keyword matching, hindering automation. To tackle this, the work proposes the first system integrating Retrieval-Augmented Generation (RAG) with a multi-agent collaboration framework. The system orchestrates evidence retrieval, evaluation generation, adversarial challenge of assertions, and resolution of disagreements to produce interpretable audit recommendations that include citations, reasoning, gap analysis, and remediation guidance. Experimental validation under the ISO/IEC 27001 standard demonstrates that the approach effectively supports control interpretation and audit preparation, significantly improving efficiency. Nevertheless, human oversight remains necessary to calibrate judgments of evidentiary sufficiency.
This study addresses the lack of rigorous statistical guarantees in sequential sampling for auditing by formulating it as a sequential hypothesis test under sampling without replacement from a finite population. It defines null and alternative hypotheses based on a tolerable deviation rate and constructs exact stopping and decision rules that provide a priori control over both Type I and Type II error probabilities. The work introduces the first sequential audit sampling framework supporting one-sided, two-stage, and truncated designs. Exact boundaries are derived using finite-population error probabilities and efficiently calibrated via Monte Carlo simulation under the least favorable deviation rate. This approach not only ensures pre-specified error control but also accurately estimates expected sample sizes, making it suitable for attribute sampling and tests of controls.
This study addresses the fundamental ambiguity in existing contamination audits, which often conflate truly clean data with insufficient detection power. To resolve this, the authors propose a theoretical framework based on sparse mixture models of the form \( Q_\alpha = (1-\alpha)P_0 + \alpha P_1 \), leveraging control-group estimation to assess statistical power. By integrating sample-splitting certificates with a two-stage planner, the method enables calibrated and reliable auditing. Key contributions include establishing distribution-free lower-bound certificates for contamination proportion, uncovering the failure mechanism of Gaussian budget calibration in small-sample regimes, and providing a corrective solution. Empirical results demonstrate high predictive accuracy of power curves (\( R^2 = 0.83\text{–}0.98 \)) across six channels and successfully reproduce the sensitivity ranking of injected contaminations: verbatim > paraphrase > surface. The repaired budgeting scheme proves conservatively effective, and non-rejection conclusions require joint evaluation of power, budget, and validity gating.
This work addresses the gap between strong performance of large language models on medical benchmarks and their limited capacity for defensible reasoning in real-world clinical settings using heterogeneous, longitudinal electronic health records (EHRs). To bridge this gap, the authors introduce CliniCARE-Bench—the first benchmark designed for clinical auditing—comprising 25 physician-validated scenarios and 750 real-world cases derived from MIMIC-IV. Agents must integrate structured and unstructured EHR data within a constrained tool environment to issue one of four policy-compliant rulings while fully tracing their investigative process. The framework uniquely unifies longitudinal EHR interrogation, evidence provenance, policy adherence, procedural compliance, and a calibrated abstention mechanism that distinguishes “missing data” from “medical ambiguity.” Reference rulings are established via multi-model arbitration and clinical committee calibration. Evaluation across 16 systems shows four-class accuracy of 65.3%–76.1%, yet defect-free accuracy—excluding shortcut-based responses—drops markedly by 4.8–14.8 percentage points, revealing substantial overestimation by conventional metrics.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.