🤖 AI Summary
This study addresses the deficiency of search tools in providing "no-result" signals, which causes agents to hallucinate when handling unanswerable questions. Specifically, it investigates how retrieval tool refusal responses influence frozen agent behavior. Methodologically, this work proposes the "Index Void" testbed, integrating BM25 retrieval, large language model (LLM) agents, and LLM-based grounding judgments to systematically quantify how different refusal phrasings intervene in agent abstention behavior. Experimental results demonstrate that specific refusal formulations can increase the abstention rate of Qwen-series models to 97%, significantly suppressing erroneous answer generation. Furthermore, this approach outperforms conventional system prompt instructions, offering a novel paradigm for mitigating hallucinations in retrieval-augmented scenarios.
📝 Abstract
A search tool never says no: it returns its top-k passages even when the index holds no answer, so the agent sees irrelevant text instead of a miss signal. We ask what frozen search agents do when the tool refuses instead. On an index-hole testbed (257 NQ and 300 HotpotQA questions run with and without their gold passages in a 21M-passage BM25 index), seven agents receive one of five refusal wordings. An un-announced one-sentence refusal raises abstention on unanswerable questions from 23% to 97% on average for Qwen3-8B/32B and from 28% to 57% for Claude Haiku 4.5, cutting wrong answers almost one-for-one and beating a system-prompt instruction by 51 points on average. Search-R1 ignores the refusal and fabricates retrievals; Claude Sonnet 5.5 and Opus 5.5 answer from memory (abstention +2 points) and obey a system-prompt directive instead (+16). We also found that wording matters: an explanation beats a bare token; a directive inside the observation is decisive for Haiku; a soft warning is useless. Realistic triggers, from a lightweight score-based predictor to an LLM grounding judge, fall well short of the oracle, and all land on a benefit-versus-signal-quality curve that prices any trigger by its recall at a fixed false-refusal budget: for compliant agents the bottleneck is the detector inside the tool, not the agent, and the curve tells future detector work what each point of recall is worth.