Score
Designs and executes statistical audits that quantify release-side risk, detect and measure prevalence or distributional shifts, and validate outputs through empirical annotation experiments. Builds and analyzes standardized statistical reports that present metrics, uncertainty estimates, significance assessments, and interpretable recommendations for stakeholders.
This study addresses the lack of rigorous statistical guarantees in sequential sampling for auditing by formulating it as a sequential hypothesis test under sampling without replacement from a finite population. It defines null and alternative hypotheses based on a tolerable deviation rate and constructs exact stopping and decision rules that provide a priori control over both Type I and Type II error probabilities. The work introduces the first sequential audit sampling framework supporting one-sided, two-stage, and truncated designs. Exact boundaries are derived using finite-population error probabilities and efficiently calibrated via Monte Carlo simulation under the least favorable deviation rate. This approach not only ensures pre-specified error control but also accurately estimates expected sample sizes, making it suitable for attribute sampling and tests of controls.
This study addresses the lack of methodological foundations for systemic risk assessment and independent auditing under the EU’s Digital Services Act (DSA). Methodologically, it proposes the first evidence-driven compliance auditing framework, uniquely integrating legal requirements with empirical sampling theory to establish dynamic, risk-category-specific, and platform-sensitive representativeness criteria. At its core lies legally guided content sampling, augmented by stratified and temporal sampling, interdisciplinary representativeness evaluation, systemic risk classification modeling, and empirical review of compliance reporting—yielding a mixed-methods audit pathway combining qualitative and quantitative analysis. The contribution lies in overcoming key limitations of existing auditing approaches—namely, opacity and weak evidentiary chains—by empirically validating the “sampling + legal analysis” approach across diverse systemic risks, including illegal content, fundamental rights violations, democratic interference, and gender-based violence. The framework delivers an actionable, reproducible DSA compliance assessment methodology for regulators and independent auditors.
Misuse of statistical hypothesis tests severely undermines scientific reliability. This paper proposes a formal verification methodology for statistical programs: preconditions—such as normality, independence, and homoscedasticity—are explicitly encoded as logical assertions in source code; static verification is then performed on OCaml implementations using the Why3 platform to automatically detect missing or conflicting assumptions. The approach innovatively integrates contract-based programming with formal verification, distinguishing between formalizable preconditions (amenable to automated checking) and non-formalizable ones (requiring expert judgment), thereby establishing a human-in-the-loop verification paradigm. Evaluated on canonical statistical tests—including Student’s *t*-test and ANOVA—the method successfully identifies widespread misuses, such as applying the *t*-test to non-normal data or neglecting homoscedasticity checks. Results demonstrate significant improvements in the correctness, auditability, and reproducibility of statistical software.
This study addresses widespread concerns regarding methodological rigor and reproducibility in software defect prediction research, where inadequate experimental design and insufficient reporting severely undermine the credibility of findings. Conducting the first large-scale systematic audit of 101 papers published between 2019 and 2023, we employed bibliometric analysis, a structured experimental design evaluation framework, and the reproducibility assessment tool by González-Barahona and Robles to evaluate compliance with best practices in statistical methods, machine learning implementation, and result reporting. Our analysis reveals that papers exhibit an average of four methodological flaws each, with only one study fully adhering to established standards. Nearly half of the examined works omit critical details necessary for replication, and preliminary evidence suggests potential involvement of paper mill activity. These findings provide empirical grounding and actionable directions for enhancing research rigor in the field.
This study addresses the lack of a standardized evidence sampling protocol in Cybersecurity Maturity Model Certification (CMMC) assessments, which undermines the reliability and consistency of evaluation outcomes. Through an anonymous survey of 17 certified assessors combined with qualitative analysis, the research systematically reveals—for the first time—that current sampling practices are heavily reliant on assessors’ subjective experience and risk perception, with little grounding in statistical principles. The work proposes the development of a standardized, risk-based sampling framework that eschews rigid proportional rules in favor of flexible, context-sensitive guidelines. Such an approach would enhance assessment consistency while preserving essential professional judgment, thereby offering both empirical support and a novel pathway toward the formalization and improvement of the CMMC evaluation ecosystem.
Although clinical risk models often exhibit strong overall performance, they frequently demonstrate significantly disparate error rates across patient subgroups, and existing fairness audits lack reproducible stress-testing protocols. This work proposes the first systematic fairness auditing pipeline, integrating five stages: subgroup stratification, disparity measurement, mechanism diagnosis, post-hoc mitigation, and drift monitoring. The framework is rigorously evaluated on a synthetic benchmark encompassing 16 diseases, 15 social determinants of health, and predefined intersectional groups. Experiments reveal that statistical significance must be interpreted alongside minimum detectable effect sizes; calibration exhibits high variance in its impact on fairness; and mechanism diagnosis silently fails under proxy variable misspecification. The proposed group-threshold optimization consistently reduced Equal Opportunity Difference (EOD) across all 48 extrapolation scenarios, while CUSUM-based drift detection proved highly sensitive to cohort implementation, underscoring challenges in threshold transferability.
This work addresses the vulnerability of cross-site causal analysis to append-only data poisoning attacks, wherein adversaries inject plausible yet carefully crafted records to distort treatment effect estimates. The authors propose a poisoning audit framework tailored to augmented inverse probability weighting (AIPW) estimators, which precisely quantifies the worst-case causal effect bias under constraints on record plausibility, poisoning budget, and source capacity. Key contributions include a greedy scanning algorithm that efficiently computes the worst-case bias for any finite budget and sample size, and a novel Total Influence Score that unifies the direct and indirect impacts of individual records on both the propensity score and outcome models. Notably, this score yields the first conservative finite-budget bound for fully re-fitted estimators. Experiments demonstrate that the framework accurately predicts bias and that even minimal poisoning budgets can substantially compromise causal inference across multiple real-world and public datasets.
This study addresses the limited reliability of existing training data contamination detection methods in real-world auditing scenarios, particularly when distribution shifts occur or when reference benchmarks are substantially smaller than the pretraining corpus. Through a systematic evaluation of three dominant paradigms—LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC—the authors conduct 335 experiments across 27 open-source and state-of-the-art closed-source language models (up to 27B parameters). They identify distribution shift and small-scale benchmarks as two critical failure modes, revealing that only 199 evaluations yield correct conclusions. Current approaches suffer from high false-positive rates, low statistical power, or coarse-grained provenance resolution, rendering them inadequate for reliably verifying individual benchmark subsets and underscoring the irreplaceable value of transparent data provenance.
Current evaluations of AI systems predominantly rely on static benchmarks, which fail to capture behavioral risks in dynamic real-world environments. This work formalizes AI auditing as an uncertainty-aware, dynamic constraint monitoring problem across the system’s entire lifecycle, targeting critical attributes such as fairness and safety while integrating sociotechnical norms with statistical risk control. By developing a theoretical framework and supporting infrastructure for continuous auditing, the study advances AI governance beyond one-off testing toward ongoing, reliable, and accountable oversight mechanisms.