The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals

📅 2026-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently eliminating redundant logic in AI-generated code, particularly under constrained verification budgets. The authors formulate redundancy removal as a candidate statement ranking and scheduling task, proposing to control deletion order rather than relying on model confidence scores. They introduce a two-stage hybrid scheduling mechanism that combines a static shortest-first rule with a neural ranking model, prioritizing high-potential deletion candidates within limited verification resources while ensuring behavioral equivalence through execution-based validation. Experimental results on the MBPP benchmark demonstrate a 9.5% improvement in verified deletion coverage and six additional successfully solved tasks. Notably, even without in-domain validation, the static prefix strategy alone guarantees non-degrading performance.
📝 Abstract
Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software. Prompt-driven "vibe coding" is additive: new branches, guards, and fallbacks accumulate faster than obsolete logic is removed. We study the inverse problem-how an Al system should remove code when execution-verification capacity is finite. We formulate redundant-code reduction as proposal scheduling: a ranker orders single-statement deletion candidates, an execution suite accepts the first candidate that passes, and a budget bounds how many candidates may be tested. Our central observation is that candidate order, not model confidence, is the control surface a deployment can reason about. DELSCOUT instantiates two schedules. Given representative target-domain validation, a five-slot budget spends three slots on deterministic shortest-first candidates and two on complementary learned candidates; across nine MBPP replications with 0.5B, 0.6B, and 8B rankers this raises verified-deletion coverage by 9.5% relative (+6.7 accepted tasks) while consuming slightly fewer verifier calls than the matched static baseline. Without such validation the same rankers can lose coverage under shift, so we instead evaluate the complete static prefix first and append learned candidates only afterwards; for a deterministic verifier this makes coverage and character reduction non-decreasing by construction, at a measured 4.8-62.5% increase in verifier calls. MBPP+ then erases the in-domain advantage, showing that scheduling governs search while the test suite alone governs what "preserving behavior" means. The result is an auditable division of labor: models widen the search for removable code, order bounds the damage a mis-ranked proposal can do, and execution retains authority over every committed deletion.
Problem

Research questions and friction points this paper is trying to address.

code deletion
redundant code
execution verification
program maintenance
AI-assisted programming
Innovation

Methods, ideas, or system contributions that make the work stand out.

code deletion
proposal scheduling
static-first ranking
verifier budgeting
redundant code reduction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruitong Li
University of Hong Kong
Binjie Guo
Binjie Guo
PhD Candidate,Zhejiang University
Deep LearningDeep Generative ModellingNatural Language ProcessingBrain Science
A
Aisheng Mo
Zhejiang University
G
Guowei Su
Zhejiang University
H
Han Wang
Dalian University of Technology
J
Jie Li
Independent Researcher
R
Ru Zhang
Zhejiang University