perform program slicing

Designs and implements analyses and tools that compute program slices—sets of program statements and their control/data dependencies that are relevant to a given slicing criterion—and extracts, groups, or reasons about those affected code regions to decompose programs, localize changes or faults to specific slices, and reduce redundant or inconsistent modifications.

performprogramslicing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents

Jul 25, 2025
JC
Jianming Chang
🏛️ Southeast University | Singapore Management University | University of Alberta

Traditional static slicing relies on expensive dependency graph reachability analysis, limiting scalability to large or syntactically incomplete programs; learning-based approaches offer robustness but suffer from insufficient precision. This paper proposes SliceMate—the first LLM-powered multi-agent framework for static slicing—comprising synthesis, verification, and optimization agents. It eliminates dependency graph construction, enables inter-procedural and cross-file incremental scanning, and supports automatic repair. Key innovations include program dependency reasoning, dynamic scope expansion, and a joint completeness-conciseness verification mechanism, complemented by a convergence control module. Evaluated on SliceBench—a curated benchmark of 2,200 Java/Python programs—SliceMate significantly outperforms state-of-the-art tools, achieving precise slicing on programs up to 8,577 lines while unifying high accuracy with strong scalability.

Addresses scalability issues in traditional dependency graph methodsHandles incomplete syntax and large programs more effectivelyImproves static program slicing accuracy using LLM-powered agents

SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding

Sep 21, 2025
PH
Pengfei He
🏛️ University of Manitoba | Concordia University

Static slicing of incomplete or unparsable code fragments remains challenging due to inaccurate dependency modeling and hallucinated token generation in existing approaches. Method: We formulate slicing as a constrained sequence-to-sequence generation task and propose a lightweight language model–driven framework (e.g., CodeT5+), featuring: (1) a copy mechanism ensuring all output tokens are strictly sourced from the input, thereby improving dependency reasoning fidelity; and (2) dual lexical and syntactic constraint decoding, where syntactic constraints enforce TSED monotonicity to eliminate redundancy and hallucination. Results: Evaluated on CodeNet and LeetCode, our method achieves up to a 27% absolute improvement in ExactMatch over state-of-the-art methods. It demonstrates strong robustness against common real-world imperfections—including missing code segments and syntactic errors—while maintaining computational efficiency.

Existing models generate slices with extraneous tokens violating structural integrity requirementsLearning-based approaches suffer from inaccurate dependency identification between code elementsTraditional static slicing tools require complete parsable code limiting real-world applicability

Automated diagnostic support—such as error localization, proof simplification, and result preservation—is lacking in deductive verification of probabilistic programs. Method: This paper introduces the first slicing-based user diagnosis framework tailored to quantitative assertions. Its core innovations include: (i) the first formal definition of error-localizing slices; (ii) three semantically rigorous slice types—error-witness slices, refutable slices, and truth-preserving slices—formally modeled in the HeyVL language; and (iii) Brutus, a tool implementing multi-objective slice search via SMT solving, unsatisfiable core analysis, Minimal Unsatisfiable Subset (MUS) enumeration, and binary minimization. Results: Evaluated on both established and novel benchmarks, Brutus efficiently generates compact, information-rich slices; guarantees correctness of diagnostic guidance; and enables cross-language diagnostic transfer across probabilistic programming languages.

Error localization for probabilistic program verificationGenerating diagnostic slices for proof simplificationPreserving verification results with quantitative assertions

This work addresses the challenge that current language models struggle to accurately capture data dependencies in static program slicing, often generating hallucinated code. To mitigate this, the authors reformulate program slicing as a sequence-to-sequence prediction task and introduce a data-flow-aware pretraining strategy, which incorporates statement reordering based on data-flow graphs and span masking. This approach is complemented by a constrained decoding mechanism that jointly respects lexical and syntactic correctness. Evaluated on Java and Python benchmarks using compact language models such as CodeT5+, the proposed method substantially outperforms existing approaches, achieving up to a 22% improvement in ExactMatch accuracy. The results demonstrate enhanced slicing precision and effective suppression of code hallucinations.

dataflow modelingdependency modelinghallucination

Divide, Conquer and Verify: Improving Symbolic Execution Performance

Oct 05, 2023
CS
Christopher Scherb
🏛️ University of Applied Sciences and Arts, Northwestern Switzerland

Symbolic execution provides formal verification guarantees but suffers from path explosion and high SMT-solving complexity, limiting scalability to real-world software. To address this, we propose the first divide-and-conquer symbolic execution framework based on program slicing: the program is decomposed into independent slices, each executed symbolically in isolation; memory and control-flow side effects are then modeled and incrementally merged, thereby avoiding global path explosion. Our core contributions are (1) a composable side-effect merging mechanism and (2) a constraint decomposition strategy—enabling, for the first time, incremental and formally verifiable modular symbolic execution. Experimental evaluation demonstrates that our approach achieves a 3.2× speedup in path exploration while reducing memory overhead by 57%, all while preserving formal correctness guarantees.

Addressing path explosion and SMT complexity in symbolic executionImproving performance for real-world software verificationUsing divide-and-conquer with sliced execution and side effects

Latest Papers

What's happening recently
View more

This study addresses the significant challenge of verifying termination in real-world C/C++ programs, where loop interactions and nondeterministic inputs complicate analysis. The authors propose a lightweight, tool-agnostic, source-level preprocessing approach that isolates loop obligations via loop slicing and enhances termination analysis by generating input-driven concrete variants tailored to specific scenarios. An empirical evaluation integrating six termination analyzers on 117 real programs demonstrates that slicing conservatively achieves structural isolation, while concretization improves detectability in targeted scenarios at the cost of reduced semantic coverage. Crucially, the combined effect of these techniques is non-additive, indicating that preprocessing should complement—rather than replace—analysis of the original program. The work further reveals substantial variation in how different analyzers respond to preprocessing, offering practical guidance for adaptive usage by developers.

loop interactionsnon-terminationnondeterministic inputs

This work addresses the limited semantic precision of large language model (LLM) agents in function- and line-level fault localization, which stems from existing graph-based program understanding approaches’ inadequate modeling of intra-procedural value flows. To overcome this, the authors propose a multi-granularity program dependence graph that refines nodes to the statement level and explicitly captures data flow through definition-use edges. For the first time, intra-procedural dataflow slices are integrated as first-class primitives into a repository-scale graph representation. A framework-agnostic, three-layer API toolkit is also introduced, enabling LLMs to directly query these slices for accurate error localization and patch generation. Evaluated on SWE-bench Lite, the approach improves Function Recall@1 by 17.0 points, Line Recall@1 by 15.0 points, and achieves a repair success rate (Pass@1) of 22.0%, outperforming the baseline by 4.7 percentage points.

automated program repairdata-flow analysisfault localization

Program analyses often lack robustness in the face of code changes. This work introduces, for the first time, a unified framework grounded in category theory that formalizes programs and their properties as categorical objects, capturing various forms of robustness—such as variable renaming and semantic refinement—via structure-preserving functors. Two implementation pathways are proposed: one lifts constructions from restricted computational models to general-purpose programs, while the other ensures stability in the composition of robust operators within algebraic program analyses. The framework not only uncovers common principles underlying loop summarization and termination analysis but also provides a theoretical foundation and predictability guarantees for developing program analyses that are more resilient to program transformations.

category theoryprogram analysisrobustness

This work addresses the inherent limitations of individual program analysis techniques—particularly their constrained precision, coverage, and insight—which hinder comprehensive software reliability assurance. Through a systematic mapping study of 248 relevant publications, the paper presents the first taxonomy of combined program analysis approaches explicitly centered on synergistic effects and interaction patterns. The proposed multidimensional classification framework is structured around three core dimensions: collaboration objectives, workflow architectures, and types of mapping functions. This framework systematically uncovers commonalities and distinctions in the design of existing methods, offering a clear conceptual foundation for understanding, comparing, and developing novel combined analysis techniques. Furthermore, it delineates current research trends and identifies promising directions for future investigation.

combined techniquesprogram analysissoftware dependability

Hot Scholars

MP

Michael Pradel

Faculty, CISPA Helmholtz Center for Information Security • Professor, University of Stuttgart
Software EngineeringProgramming Languages
XP

Xin Peng

East China University of Science and Technology
Artificial IntelligenceMachine LearningComplex Process Modeling
YL

Yuekang Li

Lecturer (Assistant Professor), University of New South Wales
Software EngineeringSoftware SecurityAI Red Teaming
MK

Miryung Kim

Professor and Vice Chair of Graduate Studies, UCLA Computer Science
Software Engineering
BR

Baishakhi Ray

Associate Professor, Columbia University
Software EngineeringMachine LearningAI4CodeAI4SE