failure-aware retry

Designs and builds runtime mechanisms that detect failures during execution and identify pivotal decisions or states responsible for errors. Creates targeted, local retry and exploration policies that steer new attempts away from past mistakes to recover at test time and complete tasks autonomously.

failure-awareretry

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$224K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited generalizability of existing fault management approaches, which rely on task-specific multimodal processing pipelines. The authors propose RuntimeSlicer, the first framework to achieve a unified, cross-modal, and task-agnostic runtime system state representation. By integrating unified runtime contrastive learning with temporal consistency modeling, RuntimeSlicer aligns and encodes metrics, traces, and logs into a shared system state embedding space. It further enables lightweight adaptation to diverse downstream tasks through unsupervised state segmentation and state-aware fine-tuning. Experimental evaluation on the AIOps 2022 dataset demonstrates that RuntimeSlicer achieves strong generalization capability and practical feasibility in both system state modeling and fault management tasks.

failure managementgeneralizationmultimodal data

This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.

agent failuresdiagnostic signalsfailure modes

Existing automated program repair techniques struggle to address complex logical errors and silent failures due to their inability to accurately model runtime dynamic behaviors and data dependencies. This work proposes TraceRepair, a novel framework that, for the first time, incorporates runtime execution traces as shared constraints within a multi-agent collaboration mechanism. In this approach, a probe agent captures snapshots of critical variables, while multiple specialized agents—powered by large language models—perform cross-validation and iterative refinement to enable precise, dynamic-reasoning-driven repairs. Evaluated on Defects4J, TraceRepair successfully fixes 392 bugs, substantially outperforming current LLM-based methods, and demonstrates strong generalization capabilities on a newly curated dataset of recent vulnerabilities.

Automated Program RepairDynamic Data DependenciesLogic Errors

Latest Papers

What's happening recently
View more

Current LLM agents lack reproducible, intervenable, and verifiable debugging mechanisms when encountering tool failures such as timeouts, stale data, or description contamination. This work proposes the first fault reproduction–intervention–verification workflow tailored for the Model Context Protocol (MCP), implemented in an open-source web-based workbench. The platform supports real tool-call recording, injection of 12 failure types, cache-matched replay, and real-time retry capabilities. By integrating deterministic rules with an LLM-based adjudicator, the framework enables controllable behavior reproduction and rigorous performance validation. In experiments across five agents and 120 scenarios, the strongest agent completed 105 tasks; notably, the retry mechanism boosted success rates from 30% to 100% for timeout errors, while handling stale data remains challenging.

deployment reliabilityfault reproductionLLM agents

This study addresses the challenge of “silent failures” in long-running LLM agent systems, where errors are often masked as fluent and plausible yet incorrect narratives, impeding timely intervention. Through a longitudinal analysis of a personal assistant LLM agent continuously operating since March 2026 over an eight-week period, the authors conduct root-cause investigations of 22 incidents to propose the first taxonomy of five failure mechanisms specific to LLM agents and formally define the phenomenon of “fail-plausible” behavior. Leveraging a production-grade architecture—comprising 40 scheduled tasks, 8 LLM providers, tool governance agents, and a memory layer—alongside 4,286 unit tests, 827 governance checks, and manual retrospective audits, the study reveals that 70% of silent failures were only detectable by users, while retrospective auditing prevented 87% of recurrence but offered no preemptive mitigation. Failures exhibited latency up to 60 days and predominantly originated from inter-component gaps.

autonomous runtimeerror observabilityfail-plausible

This work addresses the challenge of robotic policy degradation in real-world environments due to repeated failures, a problem often mitigated by manual intervention in existing approaches. The authors propose the Failure-Aware Retry (FAR) framework, which, during deployment, leverages observed failure cases to construct contrastive preference signals that guide the policy away from ineffective actions. FAR integrates lightweight action perturbations to encourage local exploration and employs online replay of successfully recovered trajectories for continual policy refinement. Notably, the method enables autonomous recovery without additional training, substantially improving both data efficiency and robustness. Experimental results demonstrate that FAR increases average task success rates by 17.6% in simulation and 11.7% on physical robots, with particularly strong performance under constraints on reset counts and time steps.

autonomous retrypolicy improvementreal-world deployment

This work addresses the challenge of determining whether a local recovery point is semantically valid when structured tool-using agents fail mid-execution, particularly in scenarios where downstream components have already committed to outputs from upstream stages. The paper introduces DART, a runtime system that formalizes the notion of “semantic recoverability” for the first time. DART enables safe and efficient partial recovery by identifying failure instances, verifying semantic boundaries, aligning checkpoints, and selecting legitimate recovery points under dependency and effect constraints. Its modular architecture incorporates explicit acceptability checks to prevent invalidation of already-committed downstream work. Empirical evaluation across three LLM-driven tasks and the LangGraph framework demonstrates that DART successfully recovers all commitment-sensitive cases where baseline methods fail, with no unsafe rollbacks detected in a five-domain safety audit.

commitment-sensitivelocal recoveryruntime failure

This study addresses the critical issue of infinite agent loops (IAL) in large language model (LLM) agents, which can arise from missing or ineffective termination conditions during iterative execution, leading to resource exhaustion and denial-of-service. The work presents IAL-Scan, the first systematic approach to detecting IAL vulnerabilities, leveraging a unified intermediate representation to abstract heterogeneous agent code, constructing agent loop dependency graphs, and applying path reachability analysis to identify high-risk feedback loops. Evaluated on 6,549 open-source projects, IAL-Scan identified 74 potential IAL instances; manual validation confirmed 68 true positives across 47 projects, achieving a precision of 91.9%. This demonstrates IAL-Scan’s effectiveness in enabling cross-framework, high-precision discovery of IAL vulnerabilities.

agent executionfeedback pathsInfinite Agentic Loops

Hot Scholars

SX

Shiyun Xiong

Harbin Institute of Technology
Natural Language Processing
ZG

Zhifeng Gao

DP Technology
Data MiningMachine LearningAI for ScienceAI for Industry
LO

Litu Ou

University of Edinburgh
Natural Language ProcessingMachine LearningInformation Retrieval
XW

Xin Wang

China Agricultural University
MechatronicsAutomationSensorsRobotics