SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding

📅 2025-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Static slicing of incomplete or unparsable code fragments remains challenging due to inaccurate dependency modeling and hallucinated token generation in existing approaches. Method: We formulate slicing as a constrained sequence-to-sequence generation task and propose a lightweight language model–driven framework (e.g., CodeT5+), featuring: (1) a copy mechanism ensuring all output tokens are strictly sourced from the input, thereby improving dependency reasoning fidelity; and (2) dual lexical and syntactic constraint decoding, where syntactic constraints enforce TSED monotonicity to eliminate redundancy and hallucination. Results: Evaluated on CodeNet and LeetCode, our method achieves up to a 27% absolute improvement in ExactMatch over state-of-the-art methods. It demonstrates strong robustness against common real-world imperfections—including missing code segments and syntactic errors—while maintaining computational efficiency.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Static program slicing is a fundamental technique in software engineering. Traditional static slicing tools rely on parsing complete source code, which limits their applicability to real-world scenarios where code snippets are incomplete or unparsable. While recent research developed learning-based approaches to predict slices, they face critical challenges: (1) Inaccurate dependency identification, where models fail to precisely capture data and control dependencies between code elements; and (2) Unconstrained generation, where models produce slices with extraneous or hallucinated tokens not present in the input, violating the structural integrity of slices. To address these challenges, we propose ourtool, a novel slicing framework that reformulates static program slicing as a sequence-to-sequence task using lightweight language models (e.g., CodeT5+). Our approach incorporates two key innovations. First, we introduce a copy mechanism that enables the model to more accurately capture inter-element dependencies and directly copy relevant tokens from the input, improving both dependency reasoning and generation constraint. Second, we design a constrained decoding process with (a) lexical constraint, restricting outputs to input tokens only, and (b) syntactic constraint, leveraging Tree Similarity of Edit Distance (TSED) monotonicity to detect structurally invalid outputs and discard them. We evaluate ourtool on CodeNet and LeetCode datasets and show it consistently outperforms state-of-the-art baselines, improving ExactMatch scores by up to 27%. Furthermore, ourtool demonstrates strong performance on incomplete code, highlighting its robustness and practical utility in real-world development environments.
Problem

Research questions and friction points this paper is trying to address.

Traditional static slicing tools require complete parsable code limiting real-world applicability
Learning-based approaches suffer from inaccurate dependency identification between code elements
Existing models generate slices with extraneous tokens violating structural integrity requirements
Innovation

Methods, ideas, or system contributions that make the work stand out.

Copy mechanism for dependency identification
Constrained decoding with lexical restrictions
Syntactic constraint using TSED monotonicity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Pengfei He
Pengfei He
PhD student, Michigan State University
Machine LearningTrustworthyAIHigh-dimensional StatisticsCausal Mediation Analysis
S
Shaowei Wang
University of Manitoba, Winnipeg, Canada
T
Tse-Hsun Chen
Concordia University, Montreal, Canada