Score
Designs, implements, and evaluates mechanisms for detecting, propagating, classifying, and recovering from runtime errors and exceptional conditions across software components. This includes building exception-handling constructs and frameworks, applying error‑handling and exception‑handling patterns, implementing retry and backoff strategies, and mapping/internalizing errors for graceful degradation and API error responses to improve robustness and maintainability.
This work addresses the challenge that large language models often generate inaccurate or incomplete exception-handling code in warehouse-scale software due to insufficient awareness of contextual and dependency information. To this end, the authors propose CatchAll, a novel approach that systematically integrates API-level exception knowledge, intra-repository calling context, and cross-repository reusable patterns to construct structured prompts that guide large language models toward generating precise exception-handling code. CatchAll achieves multi-granular knowledge injection through API–exception mapping, call trace modeling, and cross-project pattern mining. Evaluated on the newly curated benchmark RepoExEval, CatchAll substantially outperforms existing methods, achieving a CodeBLEU score of 0.31, an intent prediction accuracy of 60.1%, and a Pass@1 rate of 29%.
本文提出使用EXCODER结合大语言模型自动生成缺失的异常相关代码,以支持异常行为测试,并通过GitHub Java项目验证其有效性。
Fault propagation paths in cloud services are difficult to trace due to error-wrapping, and existing approaches suffer from insufficient accuracy. This paper proposes an iterative backward search method that synergistically integrates static analysis with a large language model (LLM) agent: it first constructs a function call graph, then leverages the LLM to semantically model logs and match candidate functions, enabling multi-step reasoning to reconstruct the full propagation path from terminal log entries to root-cause faults. To our knowledge, this is the first work to jointly exploit static code structure and LLM-based semantic understanding for fault-chain reconstruction. Evaluated on 67 production microservices and 102 real-world failures at ByteDance, our method achieves a 97.0% path reconstruction accuracy—substantially outperforming both pure static analysis and state-of-the-art LLM-based baselines.
This study addresses the lack of systematic understanding regarding test coverage of anomalous behaviors in real-world systems, particularly anomalies that do not propagate to test failures. For the first time, it jointly examines both propagated and non-propagated exceptions by dynamically instrumenting 25 Python projects, monitoring 5,372 methods, 17.9 million method calls, and 1.4 million exceptions. The analysis reveals that 21.4% of methods raise exceptions, with approximately 20% of these doing so frequently—exhibiting a median rate of one exception per ten invocations. These findings challenge the conventional assumption that exceptions are rare events and demonstrate that anomalous behavior is far more prevalent in practice than previously believed.
Large language models (LLMs) frequently generate code lacking robust exception handling, leading to runtime fragility. To address this, we propose Seeker—a novel multi-agent framework that systematically orchestrates LLMs across the full exception-handling pipeline: detection, exception type identification, and repair generation. Seeker comprises five specialized agents—Scanner, Detector, Predator, Ranker, and Handler—that jointly integrate static analysis, exception pattern mining, and ranking-enhanced repair generation. This design mitigates three critical challenges: inaccurate fragile-code identification, erroneous exception-type classification, and semantically distorted repairs. Evaluated on multiple open-source projects, Seeker achieves a 37.2% improvement in exception-handling coverage, an 89.5% accuracy in exception-type identification, and an over-82% compilation success rate for generated fixes—demonstrating significant advances in automated, LLM-driven exception handling.
This study addresses the persistent occurrence of software defects after release, particularly in C/C++ and Java systems, whose underlying causes remain poorly understood. Through a large-scale empirical analysis of over 14,000 open-source projects, the work systematically compares pre-release and post-release defect characteristics using multidimensional metrics—including code complexity, size, change frequency, and development history—and employs statistical modeling to uncover key patterns. It reveals for the first time that post-release defects are significantly concentrated in legacy modules that undergo frequent modifications, with their root causes primarily stemming from dynamic evolutionary pressures rather than static code structure. Furthermore, such defects exhibit longer repair cycles and higher complexity, offering empirical grounding for targeted testing strategies and improved reliability assurance.
This work addresses a critical yet overlooked reliability issue in code generated by large language models (LLMs): despite passing compilation and unit tests, such code often fails in deployment due to structural inconsistencies—such as missing configurations, invalid imports, or omitted security controls—that evade detection by conventional CI/SAST tools. The paper introduces the “patchwork problem” to characterize these cross-module global defects, proposes an eight-category taxonomy specific to LLM-generated code, and formalizes structural consistency via invariants derived from a multidimensional code graph encompassing imports, calls, dependencies, configurations, and routing. Building on this foundation, the authors design a hybrid verification framework that integrates traditional static analysis with custom graph-based invariant checkers to precisely identify structural flaws invisible to existing tools. Empirical evaluation reveals that such defects are pervasive across major LLMs under diverse prompting strategies and exhibit distinct model-specific patterns.
This study addresses previously unexamined runtime failures in Model Context Protocol (MCP) servers, such as accepted configuration parameters that are not enforced, leading to unintended default behaviors and system unreliability. Through manual analysis of 837 runtime failure reports from 473 active MCP repositories, the authors employ a bottom-up open coding approach to construct the first comprehensive taxonomy of MCP runtime failures, encompassing dimensions such as protocol interactions, tool invocations, and state management. The resulting classification comprises 11 top-level categories and 27 subcategories, totaling 73 leaf-level failure types. Validated by 55 developers—each encountering an average of 20 categories—and supported by empirical observations across all categories, the taxonomy demonstrates broad applicability and strong external validity.
This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.
This study addresses the absence of real-world repository benchmarks and execution security risks in using large language models (LLMs) to repair runtime errors. It introduces HealBench, the first benchmark for runtime crash repair over real codebases, alongside HealGuard, a trusted repair protection framework based on taint analysis. By integrating static and dynamic taint analysis with HealCore—a restricted Python subset—the proposed approach constructs an automated repair agent. Experimental results demonstrate that under optimal configurations, the method achieves a 38.11% program execution recovery rate and a 28.68% test pass rate while successfully identifying 17.4% of potentially unsafe repairs. These findings validate both the effectiveness and security of LLM-driven automated crash repair in real-world scenarios.