Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of real-world repository benchmarks and execution security risks in using large language models (LLMs) to repair runtime errors. It introduces HealBench, the first benchmark for runtime crash repair over real codebases, alongside HealGuard, a trusted repair protection framework based on taint analysis. By integrating static and dynamic taint analysis with HealCore—a restricted Python subset—the proposed approach constructs an automated repair agent. Experimental results demonstrate that under optimal configurations, the method achieves a 38.11% program execution recovery rate and a 28.68% test pass rate while successfully identifying 17.4% of potentially unsafe repairs. These findings validate both the effectiveness and security of LLM-driven automated crash repair in real-world scenarios.
📝 Abstract
Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, each paired with a reference execution on the patched version. HealBench also provides a unified framework that lets LLM agents heal with cross-file context and live runtime state. We then design HealGuard, which requires healing code to be written in HealCore, an analyzable subset of Python, and uses static and dynamic taint analysis to check whether state changed by healing reaches operations protected by developers. We evaluate a dedicated healing method and three general coding agents with three backbone LLMs. The best setting resumes execution in 38.11% of instances and passes the target test in 28.68%, showing that existing agents can already heal a meaningful share of real repository-level crashes. However, among executions that pass, HealGuard flags 17.4% whose healing-changed state may reach a protected operation. On 684 controlled cases, HealGuard detects all unsafe cases, at the cost of a 68.42% false positive rate.
Problem

Research questions and friction points this paper is trying to address.

runtime error healing
real-world repositories
safety concerns
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Runtime Error Healing
Benchmark
Safety Guardrail
Taint Analysis
Large Language Models