Who Broke Me? Execution-Guided Repair of Behavioral Dependency Breaks

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of behavioral breaking changes (BBCs) induced by dependency upgrades, which often occur without interface modifications and remain difficult for existing methods to accurately localize and repair. To tackle this, we propose BBCFixer, a framework that incorporates an execution-guided root cause localization mechanism to identify offending APIs by contrasting dynamic execution results across library versions. Furthermore, it integrates code-difference filtering with large language model (LLM) agents to automatically generate repair patches. As a key contribution, we construct BBCBench, the first benchmark dataset dedicated to BBCs. Experimental evaluations demonstrate that our approach increases test pass rates by an average of 17% while reducing the required repair steps by 27%, highlighting its effectiveness in automating the resolution of dependency-induced behavioral regressions.
📝 Abstract
Dependency upgrades can break downstream projects without changing the library interface. Such behavioral breaking changes are difficult for developers to fix, because the failing test does not always point to the root API, the upstream API that causes the break. Existing LLM-based repair methods obtain evidence for the repair from compiler feedback or library documentation. However, a behavioral break produces no compiler feedback and is often undocumented. Agents that receive no upgrade evidence also usually do not retrieve upstream evidence themselves, and most of their failed repairs do not identify the root API. The library diff provides useful upstream evidence, but the root API must first be identified to select the relevant part of the diff. We present BBCFixer, a repair method that runs the failing test under the old and new library versions, ranks the calls whose return value differs to identify a candidate root API, and filters the library diff with that candidate. To evaluate repairs of behavioral breaks, we further introduce BBCBench, a benchmark of 100 behavioral dependency breaks in Python and JavaScript. We compare BBCFixer on BBCBench with two baselines: (1) Pure Agent, an agent that receives no upgrade evidence, and (2) BDUpdater, a method that mines library documentation. BBCFixer raises the average pass rate by up to 17% relative to Pure Agent and needs up to 27% fewer agent steps per successful repair. BBCFixer thus provides a way to select the upstream evidence behind a behavioral break, so that an LLM agent can repair behavioral breaking changes more effectively and efficiently.
Problem

Research questions and friction points this paper is trying to address.

behavioral breaking changes
dependency upgrade
automated program repair
root API identification
LLM-based agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavioral Dependency Breaks
Execution-Guided Repair
Root API Identification
Library Diff Filtering
BBCBench
🔎 Similar Papers