How Should Diffusion Language Models Edit Code?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between localization prediction and content generation capabilities in diffusion language models for code editing. Building upon masked diffusion language models, this work systematically evaluates four interaction paradigms—including whole-file rewriting and search-and-replace—using the CanItEdit benchmark alongside a Wiki probing framework. Furthermore, it introduces a "composition gap" metric to quantify the degree of separation between localization and generation. The findings reveal that strong infilling proficiency alone is insufficient to guarantee editing reliability. Instead, dependable code editing necessitates simultaneously achieving complete change coverage and minimizing extraneous regeneration. These insights provide a theoretical foundation for designing code editing interfaces that effectively balance precision with safety.
📝 Abstract
Code editing requires a model to decide where to make changes, generate the new content, and preserve everything else. We study how masked diffusion language models divide these responsibilities across four editing interfaces: whole-file rewriting, search-and-replace, locate-then-infill, and token-level editing. Experiments on CanItEdit reveal a composition gap: diffusion models can generate coordinated changes when the correct edit locations are supplied, but much of this capability is lost when those locations must be predicted. Access to the intact original code helps the model fill multiple edit regions, yet does not resolve the difficulty of selecting those regions. By varying the editable regions while holding the generation model and decoding procedure fixed, we identify two distinct requirements for successful editing: covering every required change and placing precise boundaries around it. Missing a required region prevents the corresponding change, while widening regions to ensure coverage can sharply reduce success by requiring unchanged code to be regenerated. A sentence-level Wiki editing probe shows the same qualitative gap between supplied and predicted locations beyond code. These findings show why strong infilling capability alone does not ensure reliable editing: the interface must expose all required changes while limiting regeneration of unchanged code.
Problem

Research questions and friction points this paper is trying to address.

diffusion language models
code editing
composition gap
edit localization
infilling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Language Models
Code Editing
Composition Gap
Infilling
Editing Interfaces
🔎 Similar Papers
No similar papers found.