code reading and debugging

Reads and analyzes source code to understand program behavior, execution flow, and data/state transformations in order to locate and diagnose defects. Designs and applies debugging artifacts and processes—such as breakpoints, stack-trace analysis, log inspection, minimal reproducers, and targeted test cases—to isolate root causes and validate fixes.

codereadinganddebugging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.

debuggingexecution tracehypothesis generation

How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach

Feb 25, 2025
MS
Marius Smytzek
🏛️ CISPA Helmholtz Center for Information Security | Humboldt-Universität zu Berlin

This paper addresses the dual challenges of low fault localization accuracy and weak root-cause interpretability in software debugging. To this end, we propose an interpretable diagnosis method based on multi-execution feature fusion. Through empirical analysis of 310 real-world defects, we first establish—systematically and for the first time—that scalar pairs constitute the strongest failure-correlated features. Building upon this insight, we design a joint modeling framework that integrates 17 fine-grained execution features, including variable values, branch conditions, and definition-use chains. We further develop a feature-importance-driven interpretable decision tree model that automatically generates human-readable diagnostic rules. Evaluation across 20 open-source projects demonstrates that our approach significantly improves both fault localization accuracy and root-cause identification depth, substantially reducing developer debugging time. The method achieves a favorable balance between high precision and strong interpretability.

Analyzing diverse execution featuresDeveloping interpretable debugging diagnosesEnhancing fault localization accuracy

InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration

Oct 21, 2025
YW
Yunkun Wang
🏛️ Zhejiang University | Alibaba Group

LLMs frequently generate code containing subtle, hard-to-diagnose logical errors; existing self-repair methods rely on static analysis or shallow execution logs, lacking human-like interactive, dynamic debugging capabilities. This paper introduces InspectCoder—the first intelligent code repair system enabling LLMs to perform breakpoint setting, runtime state inspection, and incremental experimentation via real debuggers, shifting the paradigm from trial-and-error to root-cause localization. Its key contributions are: (1) a novel LLM-debugger collaborative agent framework supporting stateful, adaptive dynamic analysis; (2) real-time debugging feedback integrated as a process reward to guide multi-step reasoning optimization; and (3) seamless integration with the open-source middleware InspectWare, ensuring compatibility with mainstream Python testing frameworks. On BigCodeBench-R and LiveCodeBench-R, InspectCoder achieves 5.10–60.37 percentage points higher repair accuracy than the strongest baseline, while improving efficiency by 1.67×–2.24×.

Existing methods lack interactive dynamic analysis capabilitiesInspectCoder enables LLMs to conduct active debugging via debugger controlLLMs generate buggy code with complex logic errors

Automated Defects Detection and Fix in Logging Statement

Aug 06, 2024
RZ
Renyi Zhong
🏛️ The Chinese University of Hong Kong | Singapore Management University

Low-quality log statements—such as ambiguous or misleading ones—obscure actual program behavior and impede software maintenance. Prior work primarily focuses on detecting single log defects and relies on manual fixes. This paper proposes LogFixer, the first automated two-stage framework targeting four real-world log defects: detection and repair. In the offline stage, a lightweight similarity classifier is trained on synthetically defective logs; in the online stage, problematic logs are identified via joint modeling of static textual features and dynamic variable contexts, and semantically appropriate repairs are recommended using large language models (LLMs). LogFixer innovatively integrates a lightweight classifier with LLMs in a synergistic paradigm, ensuring robust detection while enhancing repair validity. Evaluation shows an F1-score of 0.625; adoption rates of static and dynamic repair suggestions improve by 48.12% and 24.90%, respectively; repair suggestion adoption reaches 61.49% on unseen projects; and 40 fixes submitted to GitHub have yielded 25 merged confirmations.

Detects and repairs defects in logging statements automaticallyIdentifies four types of logging defects via log-centric analysisImproves log quality using LLM-based recommendations and fixes

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing large language model–based automated program repair approaches, which rely on end-to-end test feedback and struggle to precisely identify internal logical deviations. To overcome this, the authors propose SpecTune, a framework that inserts checkpoints along execution paths to generate localized postconditions and evaluates intermediate program behaviors against dynamic execution results, thereby providing fine-grained debugging signals. SpecTune introduces an intermediate behavior reasoning mechanism and designs two key signals—a specification validation signal (α) and a discriminative signal (β)—to substantially enhance the reliability of automatically generated specifications and the precision of repairs. Experimental results demonstrate that SpecTune significantly outperforms current baseline methods in both fault localization accuracy and repair success rate.

Automated Program RepairFault LocalizationIntermediate Behavioral Signals

This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.

defect characterizationpost-release defectsresidual faults

Debugging in data-intensive programming faces significant challenges, including fragmented evidence, difficulty in discerning discrepancies between expected and observed behaviors, and the complexity of tracking state evolution across components. Through semi-structured interviews and thematic analysis, this study systematically characterizes practitioners’ debugging practices and, for the first time, identifies three core requirements: cross-artifact evidence alignment, expectation-based comparison mechanisms, and traceable state evolution. Building on these insights, the work constructs a visualization-driven design space tailored to debugging in data-intensive contexts, exposing critical gaps in existing tools and providing a theoretical foundation and clear direction for the development of future debugging aids.

data-intensive programmingdebugging challengesevidence-driven reasoning

Bug fixing is a complex and time-consuming task in software development. Bug localization research tends to focus on the accuracy of automated tools that suggest source code files for developers to look at. However, little is known about how developers use these tools in practice. This paper reports on an ongoing qualitative user study. Eleven participants worked through four realistic bug localization tasks in a controlled environment and were given varying levels of support information offered by a specialized tool. Participants were asked to think aloud in a semi-structured interview session. The preliminary findings provide insight into three aspects of practice: how developers interact with tools, the role social and contextual information plays, and problem solving. The study demonstrates that bug localization is complex and suggests that the adoption of effective tools depends on more than their accuracy.

bug localizationdeveloper behaviourqualitative study

Hot Scholars

CM

Collin McMillan

University of Notre Dame
Automated Software EngineeringAI4SEEye-TrackingHuman Attention
DG

Debin Gao

Singapore Management University
computer security
NK

Natalie Kiesler

Professor for Teaching and Learning in Higher Education, Computer Science, Nuremberg Tech
Computing EducationCompetencyAI FeedbackOpen Science
TS

Tania Stathaki

Imperial College London
Object TrackingImage FusionImage RegistrationImage Processing
RH

Ruidong Han

Meituan
recommender systemgenerative model