perform program comprehension

Designs and implements analyses and tools that reconstruct source-level structure and intent from static artifacts and runtime information, including resolving symbols and mapping execution traces back to code. Builds methods to identify missing identifiers or imports, recreate execution contexts and minimal scaffolding, and bind and execute candidate functions so that program behavior and observability signals can be understood and manipulated.

performprogramcomprehension

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Reimagining Disassembly Interfaces with Visualization: Combining Instruction Tracing and Control Flow with DisViz

Oct 21, 2025
SH
Shadmaan Hye
🏛️ SCI Institute | Lawrence Livermore National Laboratory

Binary disassembly analysis suffers from ambiguous source-to-instruction mapping and difficulty in jointly preserving execution order and control flow. To address this, we propose DisViz—a performance-analysis-oriented, interactive disassembly visualization tool. Its core contributions are threefold: (1) a basic-block–based instruction layout that explicitly preserves execution order while intuitively revealing control structures (e.g., loops); (2) block-level minimaps to enhance contextual awareness and navigation in large-scale disassembly; and (3) integrated instruction tracing, control-flow graph visualization, and dynamic source-code correlation, enabling bidirectional, web-based navigation between source and disassembly. An empirical evaluation with ten domain experts from diverse institutions demonstrates that DisViz significantly improves both accuracy in identifying compiler optimization behaviors and overall analysis efficiency—validating its effectiveness for understanding compilation transformations and their performance implications.

Addressing challenges in mapping binary instructions to source codeImproving developer comprehension of compiler optimizations in binariesVisualizing disassembly code with execution order and control flow

reAnalyst: Scalable Annotation of Reverse Engineering Activities

Jun 06, 2024
TZ
Tab Zhang
🏛️ Ghent University | Lawrence Livermore National Laboratory | The University of Arizona

Traditional reverse engineering (RE) research relies on manual data collection and subjective analysis, suffering from low efficiency, poor scalability, and insufficient objectivity. To address these limitations, this paper proposes reAnalyst—a tool-agnostic, multimodal RE behavior acquisition and semi-automatic annotation framework. reAnalyst synchronously captures heterogeneous data—including screen screenshots, keyboard inputs, and process information—and leverages computer vision–driven activity recognition, heuristic behavioral modeling, and semi-supervised learning to accurately identify and semantically annotate RE operations. Experimental results demonstrate high recognition accuracy across diverse and complex screenshots. Empirical evaluation confirms the framework’s effectiveness and practical utility, with broad endorsement from professional reverse engineers. This work significantly advances the automation, reproducibility, and large-scale analytical capability of RE research.

Enables efficient analysis of protection techniquesFacilitates study of reverse engineering practicesOvercomes manual data collection limitations

This work addresses the limitation of large language models (LLMs), which, trained solely on static code, lack the deep, long-horizon reasoning capabilities essential for real-world software development. To bridge this gap, the authors propose a novel “understanding through refactoring” paradigm that reconceptualizes the development process as a refactorable multi-agent trajectory. By inversely synthesizing high-quality reasoning trajectories—encompassing planning, debugging, and iterative refinement—from static code, and integrating dependency graph–guided trajectory generation with search-based chain-of-thought optimization, the method enables continuous pretraining. Experiments on Llama-3-8B demonstrate significant improvements in long-context comprehension, programming proficiency, and agent-like behavioral capabilities, effectively enhancing the model’s capacity for deep reasoning.

code generationLarge Language Modelslong-horizon reasoning

Combining Static Analysis Techniques for Program Comprehension Using Slicito

Mar 19, 2025
JK
Jan Kofrovn
🏛️ Charles University

Existing program comprehension tools struggle to balance scalability and precision in static analysis. This paper addresses C# programs by proposing an interactive, progressive analysis framework: developers first employ lightweight interprocedural data-flow analysis to rapidly identify critical code subregions; subsequently, high-precision symbolic execution is selectively applied to those regions. The framework introduces a novel composable analysis and visualization architecture—inspired by Moldable Development—that enables on-demand assembly of customized comprehension tools directly within Visual Studio. Evaluated on real-world industrial case studies, the approach maintains analytical efficiency while significantly improving precision, thereby enhancing reasoning about complex code behaviors. Key contributions include (1) a progressive, developer-guided analysis paradigm that bridges coarse-grained scalability and fine-grained accuracy; (2) a modular, extensible architecture supporting tool composition without recompilation; and (3) empirical validation demonstrating substantial precision gains—up to 3.2× improvement in path-sensitive defect detection—without compromising analysis throughput.

Enables interactive code scope reduction for more accurate analysis in C#.Improves program comprehension by combining scalable and precise static analysis techniques.Provides customizable analysis and visualization tools within Visual Studio.

SimpliPy: A Source-Tracking Notional Machine for Simplified Python

Oct 18, 2025
MP
Moida Praneeth Jain
🏛️ International Institute of Information Technology Hyderabad

Novice programmers often struggle due to misconceptions about Python’s control flow and scoping mechanisms. To address this, we propose a pedagogical tool integrating formal operational semantics with static program analysis. Our approach explicitly annotates operational semantics with source line numbers, enabling precise step-to-location mapping during execution; it further combines statically generated control-flow graphs with lexical scoping analysis to dynamically visualize runtime environments, call stacks, and control transfers. Implemented as an interactive, web-based debugger, the tool supports real-time exploration of program behavior. Its key innovation lies in the first unified integration of line-number-annotated operational semantics, static structural analysis, and pedagogically grounded visualization—thereby significantly enhancing beginners’ comprehension of dynamic program behavior and their efficiency in tracing execution.

Clarifying core control flow and scoping concepts for novice Python programmersHelping students build structural understanding before program tracingMaking the link between source code and execution behavior unambiguous

Latest Papers

What's happening recently
View more

This work addresses the challenging problem of recovering original source code from stripped binary functions, a task where traditional decompilation typically yields only approximate pseudocode. The paper proposes a novel paradigm that replaces pseudocode generation with direct source code retrieval. By extracting anchors such as strings and constants from binaries, the method retrieves candidate functions from a source code corpus and constructs a multimodal representation incorporating assembly instructions, decompiled code, and metadata. A large language model (LLM) is then employed for semantic re-ranking of candidates. The approach integrates Ghidra-based static analysis with an inverted index system and introduces an iterative anchor refinement strategy. Evaluated on a high-quality tcpdump dataset, it achieves 95.2% instruction coverage, and attains 35.5% coverage on general-purpose GitHub repositories, demonstrating effectiveness in both ideal and noisy real-world scenarios.

binary functionsbinary-to-source matchingreverse engineering

Static analysis struggles to reconstruct complete control flow graphs for binaries employing dynamic loading techniques—such as packed programs and modern malware—due to unresolved indirect calls. This work proposes a novel approach that integrates symbolic execution with speculative library preloading. By deploying custom hooks during symbolic execution, the method intercepts dynamic loading operations in real time, speculatively preloads required libraries, and synchronously tracks instructions while managing intercepted functions—all without executing potentially malicious code. Consequently, it safely recovers accurate control flow graphs. Experimental evaluation on 16 synthetic benchmarks demonstrates that, compared to pure static analysis, the proposed technique recovers on average 29.8% more nodes and 26.5% more edges, while achieving 100% precision and recall in library identification.

Control Flow GraphDynamic Code LoadingIndirect Calls

This work addresses a critical yet previously underexplored issue in large language model (LLM)-assisted API migration: the generation of erroneous calling contexts containing fabricated symbols—such as nonexistent imports or constructors—termed “scaffolding hallucination.” Existing evaluation metrics struggle to detect such inaccuracies effectively. To tackle this problem, the paper formally defines scaffolding hallucination and introduces a lightweight, static analysis–based fact-checking approach. The method parses the abstract syntax tree of generated code to extract referenced symbols and validates them against a knowledge base constructed from official API documentation. Evaluated on Android API migration tasks, this technique significantly outperforms conventional evaluation metrics and probabilistic judgment methods, demonstrating high precision in identifying hallucinated code while substantially reducing false positives.

API MigrationFact-CheckingHallucination

This work addresses the limitation of existing large language model–based automated program repair approaches, which rely on end-to-end test feedback and struggle to precisely identify internal logical deviations. To overcome this, the authors propose SpecTune, a framework that inserts checkpoints along execution paths to generate localized postconditions and evaluates intermediate program behaviors against dynamic execution results, thereby providing fine-grained debugging signals. SpecTune introduces an intermediate behavior reasoning mechanism and designs two key signals—a specification validation signal (α) and a discriminative signal (β)—to substantially enhance the reliability of automatically generated specifications and the precision of repairs. Experimental results demonstrate that SpecTune significantly outperforms current baseline methods in both fault localization accuracy and repair success rate.

Automated Program RepairFault LocalizationIntermediate Behavioral Signals

Hot Scholars

XR

Xiaoxue Ren

Zhejiang University
Software Engineering
MP

Massimo Poesio

Professor of Comp. Linguistics, Queen Mary University / Professor of NLP, University of Utrecht
Computational linguistics / NLPGames and NLPAnaphora / CoreferenceDisagreement and NLP
ND

Nghi D. Q. Bui

Unknown affiliation
AI4CodeSoftware EngineeringCode AgentAI4SE
NL

Nam Le Hai

Hanoi University of Science and Technology
NLPAI4CodeAI4SEContinual learning
ZL

Zhongxin Liu

Zhejiang University
Software EngineeringLarge Language Models