debugging

Designs and executes procedures, instrumentation, and experiments to find, reproduce, isolate, and eliminate defects or unexpected behavior in code, configurations, or system interactions. Builds and interprets logs, tests, breakpoints, stack traces, and diagnostic outputs and modifies code, settings, or control flow to verify fixes and prevent regressions.

debugging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.88
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach

Feb 25, 2025
MS
Marius Smytzek
🏛️ CISPA Helmholtz Center for Information Security | Humboldt-Universität zu Berlin

This paper addresses the dual challenges of low fault localization accuracy and weak root-cause interpretability in software debugging. To this end, we propose an interpretable diagnosis method based on multi-execution feature fusion. Through empirical analysis of 310 real-world defects, we first establish—systematically and for the first time—that scalar pairs constitute the strongest failure-correlated features. Building upon this insight, we design a joint modeling framework that integrates 17 fine-grained execution features, including variable values, branch conditions, and definition-use chains. We further develop a feature-importance-driven interpretable decision tree model that automatically generates human-readable diagnostic rules. Evaluation across 20 open-source projects demonstrates that our approach significantly improves both fault localization accuracy and root-cause identification depth, substantially reducing developer debugging time. The method achieves a favorable balance between high precision and strong interpretability.

Analyzing diverse execution featuresDeveloping interpretable debugging diagnosesEnhancing fault localization accuracy

Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.

debuggingexecution tracehypothesis generation

This work addresses the limitation of existing large language model–based automated program repair approaches, which rely on end-to-end test feedback and struggle to precisely identify internal logical deviations. To overcome this, the authors propose SpecTune, a framework that inserts checkpoints along execution paths to generate localized postconditions and evaluates intermediate program behaviors against dynamic execution results, thereby providing fine-grained debugging signals. SpecTune introduces an intermediate behavior reasoning mechanism and designs two key signals—a specification validation signal (α) and a discriminative signal (β)—to substantially enhance the reliability of automatically generated specifications and the precision of repairs. Experimental results demonstrate that SpecTune significantly outperforms current baseline methods in both fault localization accuracy and repair success rate.

Automated Program RepairFault LocalizationIntermediate Behavioral Signals

Who is in Charge here? Understanding How Runtime Configuration Affects Software along with Variables&Constants

Mar 31, 2025
CL
Chaopeng Luo
🏛️ National University of Defense Technology

This work reveals how dynamic interactions among program constants, environment variables, and workloads—termed PCV interactions—cause severe runtime anomalies even in configurations that pass static validation, thereby challenging conventional configuration verification paradigms. To address this, we systematically propose and empirically validate a PCV interaction model through a cross-project study involving 705 real-world configuration parameters, combining static analysis, dynamic tracing, and pattern induction to construct a case-driven risk identification framework. Our findings demonstrate that configuration exhibits a “double-edged sword” effect: its actual behavior emerges from the interplay of user intent, developer knowledge, and runtime reality. We confirm that most configurations deeply participate in runtime interactions and identify multiple high-risk interaction patterns. These results establish a theoretical foundation and practical methodology for intelligent configuration recommendation and robust configuration design.

Analyzes PCV interaction risks and patterns in large-scale systems.Examines valid configuration values causing unexpected software behavior.Studies how runtime configuration interacts with variables and constants.

This work addresses the challenge that symbolic execution engines involve numerous parameters with complex, interdependent effects, often leading users—due to limited understanding—to rely on suboptimal default configurations, while existing automated tuning approaches lack interpretability. To bridge this gap, the authors propose a human-in-the-loop parameter tuning paradigm and develop Symetra, a visual analytics system that enables dual-perspective overviews of how parameters influence branch coverage. Symetra supports interactive comparison of configuration sets and facilitates pattern recognition. Experimental results demonstrate that expert users leveraging Symetra not only accurately interpret parameter interactions and identify complementary configurations but also achieve significantly higher branch coverage and tuning efficiency compared to fully automated methods, thereby effectively overcoming the interpretability bottleneck in symbolic execution parameter optimization.

branch coverageHuman-in-the-Loopparameter tuning

Latest Papers

What's happening recently
View more

This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.

defect characterizationpost-release defectsresidual faults

Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.

behavioural gapscode coverageexpected behaviour

This study addresses the limitations of traditional statistical fault localization (SFL), which relies solely on code execution traces and often fails to accurately pinpoint root causes. To overcome this, the authors systematically incorporate execution features—such as data flow, variable values, and branch conditions—extracted via the EFDD tool from the Tests4Py dataset. They train project-specific random forest models and map feature importance back to source code lines, integrating these insights with classical SFL formulas to enhance localization accuracy. Rigorous evaluation is conducted using a confounder-adjusted mixed-effects model and paired statistical tests. Experimental results demonstrate that the proposed approach significantly improves the accuracy of reference patches while reducing inspection effort at both line and function levels, confirming its robustness and practicality across multiple dimensions.

Developer Inspection EffortExecution FeaturesFault Localization Accuracy

Hot Scholars

MR

Manuel Rigger

National University of Singapore
Software EngineeringSystemsDatabasesProgramming Languages
JS

Jin Song Dong

Professor of Computer Science, National University of Singapore
Formal MethodsTrusted AISafe AIModel Checking
TD

Tom Demeulemeester

Assistant Professor, Department of Quantitative Economics - Maastricht University
Operations ResearchCombinatorial OptimizationGame TheoryComputational Social Choice
XC

Xiaochun Cao

Sun Yat-sen University
Computer VisionArtificial IntelligenceMultimediaMachine Learning