Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

📅 2026-07-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current perturbation-based construct validity audits are highly sensitive to implementation details, yielding conclusions that lack transparency and reliability. This work proposes a self-audit framework that systematically identifies and formalizes five classes of audit failure modes (F1–F5). A case study encompassing two open-source instruction-tuned models and five safety benchmarks reveals that none of the audited units satisfy confirmatory criteria, exposing systemic vulnerabilities in prevailing practices. To address this, the paper introduces a six-point due diligence gating mechanism that establishes actionable standards for disclosing and retaining high-assurance audit evidence, thereby substantially enhancing the credibility and reproducibility of auditing outcomes.
📝 Abstract
Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported numbers. We name five classes of pipeline failure and demonstrate each in a self-audit over safety benchmarks and open-weight instruction-tuned models. Under a unified six-point due-diligence gate, every cell lands in a non-confirmatory bucket, and no cell reaches confirmatory. The evidence here is a single two-model, five-benchmark case study, and F1--F5 is an illustrative, deliberately non-exhaustive starting taxonomy -- not a comprehensive partition of audit failures. We position the gate as a withholding and disclosure protocol for assurance-grade evidence, supplementary to (not a replacement for) classical construct-validity evidence, and not as a route to benchmark-validity verdicts.
Problem

Research questions and friction points this paper is trying to address.

audit failure
benchmark validity
construct validity
AI governance
evaluation evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

audit failure modes
construct validity
due-diligence gate
benchmark auditing
AI governance
🔎 Similar Papers
No similar papers found.
Y
Yanhang Li
Northeastern University, USA
Z
Zhichao Fan
University of Illinois Urbana-Champaign, USA
Z
Zexin Zhuang
Southern Methodist University, USA