Runnable Commit Untangling for Coding Agents

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that large patches generated by coding agents are difficult to maintain, while existing splitting methods often compromise code runnability and lack empirical validation. To this end, we propose RucTangle, a novel agent-based approach that pioneers patch splitting under the strict constraint of ensuring every commit remains runnable. Furthermore, we construct TangleEval, the first evaluation framework designed to quantify the benefits of split commit histories for bug fixing. Experimental results demonstrate that RucTangle significantly reduces the non-runnable rate compared to baselines. Notably, incorporating such runnable split histories improves the Pass@1 metric of agents in fixing regression bugs by 5.2%, effectively validating the auxiliary value of high-quality commit histories for coding agents.
📝 Abstract
Coding agents produce large, tangled patches that mix multiple development purposes, making the code hard to review and maintain. Commit untangling offers the promise of organizing such large patches into untangled, manageable commits. This paper emphasizes two important limitations in existing commit untangling studies. First, they do not consider that untangled commits are ordered and should leave the code runnable. In practice, maintainers are unlikely to accept commits that prevent the code from running. Second, existing studies claim that commit untangling helps software maintenance. However, they conduct syntactic comparisons between the untangled commits and developers' original commits without directly showing the claimed maintenance benefits. To address these gaps, this paper makes two novel contributions: (1) RucTangle, the first agentic method that untangles commits while keeping the code runnable after each commit; and (2) TangleEval, the first evaluation framework that quantifies how untangled, manageable commit histories help coding agents repair bugs. We compare RucTangle against four untangling methods on 131 agent-generated patches. All histories produced by RucTangle are runnable, while baselines produce 20.6%-37.4% unrunnable commit histories. We further collect 453 agent-generated patches that introduce regressions (i.e., causing previously passing tests to fail) and ask two other coding agents to repair regressions. Augmenting agent context with RucTangle-produced histories yields 5.2% absolute improvement in pass@1. We also analyze agent trajectories to learn how they use untangled commits to navigate and fix bugs. Our findings demonstrate the value of adopting established software engineering practices in the era of coding agents, which broaden the future research agenda: how can agents actively use software history to make better development decisions?
Problem

Research questions and friction points this paper is trying to address.

Commit Untangling
Coding Agents
Runnable Commits
Software Maintenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Commit Untangling
Coding Agents
Runnable Commits
Agentic Method
Bug Repair
🔎 Similar Papers
No similar papers found.