codebase auditing

Conducts systematic inspections and analyses of software repositories and experiment code to find implementation bugs, incorrect data handling, evaluation failures (including data leakage), and sources of shortcut or spurious behavior. Produces reproducible traces and tests, quantifies the prevalence and impact of identified issues across code, and recommends concrete fixes or mitigations.

codebaseauditing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Teaching Mining Software Repositories

Jan 03, 2025
ZC
Zadia Codabux
🏛️ University of Saskatchewan | University of British Columbia | University of Florence | University of Salerno

This study addresses the lack of systematic educational resources for Mining Software Repositories (MSR) instruction targeting secondary and tertiary (undergraduate, master’s, and doctoral) students. Methodologically, it introduces the first pedagogical MSR framework by decomposing MSR methodology into teachable knowledge modules—integrating version-control analytics, defect-data extraction, communication-log mining, and qualitative-quantitative mixed-method analysis—while embedding ethical guidelines, cross-method validation, and scaffolded learning design. Its contributions include: (1) a tiered hands-on practice system; (2) instructor-facing teaching guides; and (3) ready-to-use, curated educational datasets. The resulting standardized curriculum package spans all academic levels. Empirical evaluation demonstrates that the framework significantly enhances students’ ability to conduct reproducible, ethically grounded empirical studies on real-world software repositories—thereby filling a critical gap in standardized MSR education.

Middle School StudentsMining Software Repositories (MSR) in EducationSoftware Library Education

This work addresses the limitation of existing bug reports—often lacking critical information—which hinders the effectiveness of automated program repair (APR). The authors propose TrajSpec, a novel approach that introduces trajectory-guided reasoning and hierarchical evidence representation. By leveraging execution trajectories collected via proxy runs, TrajSpec extracts multi-level evidence to refine bug reports and further enhances them through integration with repository context. Combining trajectory-guided inference, hierarchical evidence modeling, and large language model–driven report generation, TrajSpec significantly improves the Pass@1 performance of multiple APR systems on SWE-Bench Lite, raising it from 47.00% to as high as 72.00%. Ablation studies confirm the contribution of each component to the overall effectiveness.

automated program repairbug reportrepair-relevant information

The unclear spatiotemporal distribution patterns of defects in multi-fault repositories hinder cost-effective maintenance optimization. Method: We conduct an empirical study on 16 Java/Python open-source projects from Defects4J and BugsInPy, analyzing their multi-fault versions via version history tracing and precise fault localization. Contribution/Results: Our analysis reveals—temporally—that long-standing unpatched defects commonly coexist across versions, and—spatially—that defect distributions exhibit low concentration (few hotspots) and high uniformity. This challenges the conventional single-fault assumption and provides the first systematic empirical validation of widespread multi-fault coexistence and non-localized defect clustering. The findings establish a more realistic, scalable empirical foundation for test case prioritization, repair resource allocation, and evaluation of tools in multi-defect scenarios, supported by rigorously curated data.

Analyzes temporal and spatial characteristics of multi-fault systemsChallenges single-fault assumptions in Defects4J and BugsInPy datasetsInvestigates distribution and longevity of faults in Java/Python projects

Existing bug report datasets suffer from narrow coverage, poor timeliness, and incomplete metadata, hindering the application of machine learning in software quality analysis. To address these limitations, we introduce the first modern, cross-platform (GitHub/Bugzilla/Jira), cross-project (nine active open-source projects) bug report benchmark dataset, comprising over 150,000 standardized reports with comprehensive metadata and pre-split train/test splits. Our methodology includes a unified schema for semantic modeling, structured field annotation, multi-source heterogeneous data cleaning, and a Jupyter-based exploratory analysis framework. The dataset enables rigorous benchmarking for tasks including duplicate detection, RAG-enhanced generation, and automated triage—demonstrating empirically improved accuracy and relevance. Since its open release, it has become a mainstream benchmark resource for intelligent software defect analysis.

Insufficient metadata for machine learning in bug report analysisLimited scope and outdated content in existing bug report datasetsNeed for standardized datasets to support diverse SE research tasks

Analyzing Maintenance Activities of Software Libraries

Jun 09, 2023
AT
Alexandros Tsakpinis
🏛️ fortiss | Free State of Bavaria

Industrial applications heavily rely on open-source libraries, yet stalled community maintenance frequently leaves vulnerabilities unpatched for extended periods, posing critical software supply chain security risks. Existing approaches suffer from label scarcity, sparse feature representations, and incomplete modeling of transitive dependency relationships, hindering practical deployment in industrial settings. This paper proposes the first maintenance-activity monitoring framework that jointly models direct and transitive dependencies. It constructs fine-grained maintenance metrics from multi-source repository metadata—including commits, releases, issues, and pull requests—and introduces a graph propagation model to quantify the cross-dependency transmission of maintenance decay. Crucially, the method operates without manual labeling. Evaluated across multiple enterprise projects, it achieves early warning of high-risk stagnant libraries 3–6 months in advance, substantially reducing manual auditing effort and significantly enhancing the security and maintainability of open-source dependency ecosystems.

Address lack of features and labels in current researchMonitor open-source library maintenance for industrial applicationsReduce manual effort by automating dependency maintenance checks

Latest Papers

What's happening recently
View more

This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.

defect characterizationpost-release defectsresidual faults

This study addresses the challenges of high cost, error-proneness, and defect propagation in cross-repository code and test reuse during software refactoring. Through action research, the authors conduct bidirectional empirical analyses on real-world cases such as Soot/SootUp and FindBugs/SpotBugs, identifying for the first time the bidirectional reuse requirements and semantic reuse patterns inherent in refactoring scenarios. They propose a semantic alignment–based code mapping approach coupled with a hierarchical, extensible clone detection mechanism. Experimental results demonstrate that their method reduces irrelevant clones by 33%–99% on average and achieves a benchmark precision of 86%. The practical impact is further evidenced by five reported issues and ten pull requests submitted to open-source communities, eight of which have already been merged, confirming the approach’s effectiveness and applicability.

clone detectioncode reusecross-repository migration

This work addresses the significant heterogeneity in testing strategies and organization across open-source projects, which hinders rapid comprehension of their testing practices due to the lack of unified analytical tools. To bridge this gap, we propose TestMiner—a multilingual test practice analysis tool supporting languages such as Python, Java, Go, and Rust. By leveraging static code analysis and metadata extraction, TestMiner generates multidimensional visualizations encompassing test statistics, distribution, evolution, and dependency relationships. As the first tool to systematically enable cross-language and cross-ecosystem exploration of testing practices, TestMiner has been integrated into a software testing course, where it demonstrated marked effectiveness among 50 undergraduate students by significantly enhancing their understanding and critical analysis of core testing concepts, including test organization, evolution, mocking, and boundary testing.

GitHub repositoriessoftware testingtest analysis

Existing program repair benchmarks inadequately reflect real-world repository-level continuous integration (CI) scenarios, as they overlook critical challenges such as non-code artifacts, environmental dependencies, and workflow constraints. This work introduces the first repository-level repair benchmark grounded in actual GitHub Actions executions, validating patches through faithful replay of original CI workflows. The benchmark includes 567 CI failures meticulously annotated into 12 fine-grained error categories. Innovatively adopting end-to-end CI workflow re-execution as the patch validation criterion, it enables error-type-aware evaluation. By integrating log analysis, fault localization, and large language model–generated candidate patches, the approach achieves strong performance on tool-enforced errors like formatting and static checks, attaining an overall best repair success rate of 18.9%, while environment- and configuration-related issues remain notably challenging.

Automated Patch ValidationCI FailuresContinuous Integration

This study addresses growing industry concerns about the practicality and naturalness of code generated by large language models (LLMs) by systematically examining the usage patterns and defect associations of LLM-generated code and comments in active enterprise and community repositories from 2021 to 2025. For the first time, it contrasts the distribution of LLM-generated content between these two repository types through an empirical analysis integrating multiple detection tools, code clone detection, syntactic quality assessment, and manually labeled defect data. The findings reveal that the proportion of LLM-generated code has steadily declined over time and is predominantly confined to test cases, while comment generation remains stable yet exhibits low syntactic correctness. Enterprise repositories incorporate more LLM-generated content overall, which shows virtually no direct association with known defects, suggesting that such content is characterized by low risk but high functional limitations in real-world practice.

bug associationcode commentscode quality

Hot Scholars

DK

David Kohlbrenner

University of Washington
computer securitysystemscomputer architecture
TK

Tadayoshi Kohno

Professor and McDevitt Chair in Computer Science, Ethics, and Society at Georgetown University
Security and PrivacyUsable Privacy and SecuritySecurity and SocietyTech Policy
AP

Abhay Puri

Applied Research Scientist, ServiceNow Research
Agent SecurityLarge Language ModelsComputer VisionMultiModal Foundational Models
HW

Hongyi Wu

IEEE Fellow, Professor and Department Head, ECE, The University of Arizona
Intelligent and Secure Computing and Communication Systems
AN

Adam Norton

New England Robotics Validation and Experimentation (NERVE) Center, University of Massachusetts Lowell
RoboticsHuman-Robot InteractionTest MethodsInterfaces