audit statistical code

Designs and executes systematic audits of statistical analysis implementations to verify that code matches algorithmic and inferential specifications and to detect software bugs that alter estimation or hypothesis testing. This includes comparing implementation to algorithm descriptions, identifying misuse of in-sample residuals, checking that resampling methods (e.g., bootstrap) are compatible with the estimator, and quantifying how coding errors induce bias or invalid inference.

auditstatisticalcode

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the difficulty coding agents face in verifying whether programs satisfy specifications and their inability to effectively leverage expert diagnostic experience from verification failures. Building upon the executable semantics of the K framework, this work proposes a method that transforms expert diagnostics into reusable guidance, establishing a comprehensive verification pipeline encompassing specification generation, proof repair, and adequacy auditing. By designing paired clean and defective program packages, the effectiveness of the auditing mechanism is systematically evaluated. The proposed approach achieves a 164/164 pass rate on the HumanEval benchmark, with all defects accurately identified. Furthermore, experiments on KleverBench reveal optimization opportunities for guidance selection strategies under resource constraints. This research offers a novel paradigm for enhancing the formal verification capabilities of coding agents.

coding agentsformal specificationprogram correctness

StatWhy: Formal Verification Tool for Statistical Hypothesis Testing Programs

May 25, 2024
YK
Yusuke Kawamoto
🏛️ AIST | PRESTO | JST | University of Tsukuba | Kyoto University

Misuse of statistical hypothesis tests severely undermines scientific reliability. This paper proposes a formal verification methodology for statistical programs: preconditions—such as normality, independence, and homoscedasticity—are explicitly encoded as logical assertions in source code; static verification is then performed on OCaml implementations using the Why3 platform to automatically detect missing or conflicting assumptions. The approach innovatively integrates contract-based programming with formal verification, distinguishing between formalizable preconditions (amenable to automated checking) and non-formalizable ones (requiring expert judgment), thereby establishing a human-in-the-loop verification paradigm. Evaluated on canonical statistical tests—including Student’s *t*-test and ANOVA—the method successfully identifies widespread misuses, such as applying the *t*-test to non-normal data or neglecting homoscedasticity checks. Results demonstrate significant improvements in the correctness, auditability, and reproducibility of statistical software.

Automatically check requirements for statistical methods in codeFormally verify correctness of statistical hypothesis testing programsPrevent common errors in statistical program implementation

This study addresses widespread concerns regarding methodological rigor and reproducibility in software defect prediction research, where inadequate experimental design and insufficient reporting severely undermine the credibility of findings. Conducting the first large-scale systematic audit of 101 papers published between 2019 and 2023, we employed bibliometric analysis, a structured experimental design evaluation framework, and the reproducibility assessment tool by González-Barahona and Robles to evaluate compliance with best practices in statistical methods, machine learning implementation, and result reporting. Our analysis reveals that papers exhibit an average of four methodological flaws each, with only one study fully adhering to established standards. Nearly half of the examined works omit critical details necessary for replication, and preliminary evidence suggests potential involvement of paper mill activity. These findings provide empirical grounding and actionable directions for enhancing research rigor in the field.

experimental designmachine learning experimentsreproducibility

Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.

behavioural gapscode coverageexpected behaviour

This study addresses the absence of high-quality automated benchmarks for evaluating coding agents in bug verification by proposing an automated benchmark construction framework based on bug injection and behavior-preserving transformations. The approach injects defects into real-world Java projects while retaining executable witnesses, integrating dynamic execution adapters with structure-preserving transformation algorithms to enable automated verification via authentic test suites. This methodology overcomes the limitations of reusing historical bugs or relying on manually constructed data. Accordingly, a highly realistic benchmark comprising 1,300 cases is established. Experimental results demonstrate that mainstream models struggle to effectively generate executable witnesses even when defect patterns are known, revealing a fundamental technical bottleneck in current coding agents.

Automated EvaluationBenchmark ConstructionBug Validation

Latest Papers

What's happening recently
View more

Official online judging systems often misclassify logically flawed code as correct due to insufficient test suite coverage. This work presents the first systematic approach leveraging large language model–driven coding agents to automatically generate adversarial test cases that expose such testing blind spots and, in the absence of official test suites, construct effective alternatives. The study introduces a verification mechanism independent of official judges, forming a chain of evidence through multi-solution consistency checks, brute-force validation, and input legitimacy verification to ensure reliable defect detection. Evaluated on AtCoder, the method uncovered 589 misjudged submissions, with a collaborative ensemble of five agents identifying at least 906 cases. On Codeforces problems lacking official test suites, the generated test sets significantly outperformed existing baselines.

buggy submissionscode evaluationground truth

Existing large code models struggle to generate executable intermediate formal specifications, limiting precise verification and repair of program behavioral errors. This work proposes SpecCoder, a novel framework that focuses on generating executable inline assertions at critical program locations, thereby transforming static annotations into verifiable evidence. SpecCoder employs verification-guided training, fine-tuning the Qwen2.5-Coder series models using correct programs, behavioral mutants, and multi-round specification refinement trajectories. Evaluated on the HumanExec benchmark, SpecCoder substantially improves the correctness (+55.8%), completeness (+358.1%), and assertion validity (+26.6%) of inline specifications, significantly enhancing program verification and repair capabilities.

code LLMsexecutable assertionsformal specifications

This study addresses a critical gap in quantum software research: the absence of a systematic auditing mechanism for empirically grounded comparative claims, which has led to a pervasive “instantiation gap” characterized by insufficient evidentiary support. To bridge this gap, the authors propose CLAIMSTAB-QC, the first source-bound auditing framework tailored to empirical comparisons in quantum software. By integrating claim modeling, audit scope delimitation, evidence boundary identification, and directional classification, the framework enables precise validation of comparative assertions against original source materials. An evaluation across 455 claims from 119 papers reveals that only eight claims possessed sufficient matched evidence for direct auditing; among these, two were confirmed, four lacked adequate support, and two were contradicted—highlighting substantial deficiencies in the empirical rigor of current quantum software studies.

benchmarkingempirical comparisonevidence auditing

This work addresses the challenge of irreproducibility in data analysis scripts, which often stems from implicit assumptions—such as specific package versions, expected data formats, or undocumented manual interventions. The paper proposes a static analysis approach tailored to data analysis workflows that, for the first time, unifies diverse implicit assumptions into inferable constraint models. By leveraging customized program analysis and example-driven modeling, the authors develop a prototype system capable of automatically identifying these hidden assumptions, extracting executable preconditions, and generating verifiable constraints. The resulting framework supports runtime validation and automatic documentation generation, substantially enhancing script executability, reproducibility, and interpretability.

code constraintsdata analysisimplicit assumptions

Hot Scholars

MM

Maria M. Hedblom

Jönköping University
cognitive scienceapplied ontologyAIcognitive linguistics
MM

Mohamed Medhat Gaber

Adjunct Professor, Queensland University of Technology
Artificial IntelligenceData MiningData Stream MiningMachine Learning