When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing recovery mechanisms for LLM-based agents, where averaged success rates obscure the dual nature of interventions, making it difficult to distinguish beneficial rescues from detrimental disruptions. To overcome this, we formulate agent recovery as a causal decision-making problem that disentangles intervention benefits from harms by contrasting pre- and post-intervention states. We further propose a lightweight Causal Intervention Router (CIR) that leverages exclusively pre-intervention information to precisely determine optimal intervention timing, thereby preventing the disruption of inherently correct trajectories. Grounded in a causal inference framework and a dynamic policy routing algorithm, our approach demonstrates significant efficacy on long-horizon ALFWorld tasks, improving the success rate of Qwen3-14B by three percentage points while fully preserving all originally correct trajectories.
📝 Abstract
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
error recovery
causal evaluation
execution harness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Evaluation
LLM Agents
Error Recovery
Causal Intervention Router
Selective Intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuyao Xiao
School of Artificial Intelligence, Beijing Normal University
Shengling Wang
Shengling Wang
School of Artificial Intelligence, Beijing Normal University
Xuan Chen
Xuan Chen
Purdue University
AI security
K
Ke Chao
School of Artificial Intelligence, Beijing Normal University
M
Ming Cui
Ke Holdings
Feifei Qian
Feifei Qian
University of Southern California
RobophysicsLocomotionBio-inspired roboticsTerradynamics
C
Chaoyang Mei
Ke Holdings
Fanlin Meng
Fanlin Meng
Ke Holdings
Z
Ziming Yu
School of Artificial Intelligence, Beijing Normal University
J
Junxi Yin
Ke Holdings