🤖 AI Summary
This study addresses "hallucination escape" in tool-calling LLM agents, a failure mode where existing hallucination mitigation methods suppress errors under certain configurations but exacerbate them under others due to configuration conflicts. To tackle this, we propose EscapeGuard, a training-free framework that formally defines this phenomenon and reveals how conventional approaches intensify inherent conflicts. By integrating conflict-aware gating with configuration-derived attention enhancement, EscapeGuard dynamically intervenes during inference to suppress tool-selection hallucinations and prevent their cross-configuration transfer. Extensive evaluations across six benchmarks demonstrate that our approach reduces tool-selection hallucinations by 9.0 percentage points and decreases average cross-configuration hallucinations by 23.7 percentage points, achieving an 89.1% net improvement under paired-query evaluation.
📝 Abstract
Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.