metric aggregation auditing

Designs and runs audits, tests, and analysis pipelines that verify correct and consistent aggregation of metrics across code paths and experimental runs; builds detectors and statistical checks to find aggregation divergence, quantify flip rates and welfare gaps, and identify seeds or runs that cross significance thresholds.

metricaggregationauditing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.

behavioural gapscode coverageexpected behaviour

This work addresses the inefficiency of traditional fuzzing in black-box or obfuscated binary programs where static instrumentation is infeasible and control-flow feedback is unavailable. The authors propose a dynamic feedback mechanism based on Execution Divergence Graphs (EDGs), which constructs control-flow-like structures at runtime by analyzing execution traces to precisely identify path divergences and avoid redundant exploration of loops. Requiring no static program information, the approach integrates divergence detection with an EDG-guided input mutation strategy. Evaluated on multiple obfuscated targets, it substantially outperforms blind fuzzers, demonstrating its effectiveness in non-instrumented settings. Furthermore, the framework is extensible to multidimensional feedback channels, such as power consumption, broadening its applicability in side-channel-aware fuzzing scenarios.

black-box fuzzingcontrol-flow discoveryexecution traces

Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories

Jan 25, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa | Siemens AG

Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.

Data Analysis VariabilityResearch ReliabilitySoftware Engineering

Automated diagnostic support—such as error localization, proof simplification, and result preservation—is lacking in deductive verification of probabilistic programs. Method: This paper introduces the first slicing-based user diagnosis framework tailored to quantitative assertions. Its core innovations include: (i) the first formal definition of error-localizing slices; (ii) three semantically rigorous slice types—error-witness slices, refutable slices, and truth-preserving slices—formally modeled in the HeyVL language; and (iii) Brutus, a tool implementing multi-objective slice search via SMT solving, unsatisfiable core analysis, Minimal Unsatisfiable Subset (MUS) enumeration, and binary minimization. Results: Evaluated on both established and novel benchmarks, Brutus efficiently generates compact, information-rich slices; guarantees correctness of diagnostic guidance; and enables cross-language diagnostic transfer across probabilistic programming languages.

Error localization for probabilistic program verificationGenerating diagnostic slices for proof simplificationPreserving verification results with quantitative assertions

A Unit Proofing Framework for Code-level Verification: A Research Agenda

Oct 18, 2024
PC
Paschal C. Amusuo
🏛️ Purdue University | Michigan State University

Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.

Automating unit proofing to reduce manual errorsEarly detection of implementation defects in verificationEnsuring code-level correctness in large-scale software

Latest Papers

What's happening recently
View more

This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.

AI agentsfoundational assumptionsreplication

Hot Scholars

LC

Longbing Cao

Distinguished Chair Professor in AI & ARC Future Fellow (Level 3), Macquarie University
Artificial intelligenceData scienceMachine learningBehavior informatics