Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the latent failure risk in multi-agent collaboration where task outputs remain correct despite corrupted shared states. To mitigate this, we propose OffQuery, a framework that decouples the evaluation of evidence verification, state reconstruction, and task solving. We further introduce ReGround, a method designed to reconcile conflicting evidence and reconstruct trustworthy shared states. Comparative experiments across multiple foundation models (GPT, Gemini, Qwen) demonstrate that our approach yields average improvements of 309.0%, 82.9%, and 17.6% in evidence verification, state reconstruction, and task performance, respectively, within high-stakes domains such as healthcare and disaster response. These results indicate that decoupling evaluation and actively repairing shared states significantly enhances the reliability of multi-agent collaborative systems.
📝 Abstract
Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introduce OffQuery, which separately evaluates evidence verification (T1), shared-state reconstruction (T2), and task resolution (T3) in two representative high-stakes settings: healthcare and disaster response. Across GPT, Gemini, and Qwen models, standard collaboration shows much stronger task performance than state reliability. Averaged over 21 model--setting combinations, task resolution reaches 64.7%, while evidence verification and state reconstruction reach only 14.3% and 43.1%. We trace this gap to selective information use: current queries often bypass corrupted facts, which become consequential when later tasks require them. We further introduce ReGround, which resolves conflicting evidence, verifies shared facts, reconstructs a trusted state, and reasons over that state. Across seven models from three families, ReGround improves all three capabilities in every evaluated setting, with average relative gains of 309.0%, 82.9%, and 17.6% on T1, T2, and T3. Reliable collaboration therefore requires both a correct decision and a reliable shared state for future reasoning.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent collaboration
Information state corruption
Off-query failure
Shared-state reliability
Collaborative decision support
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent collaboration
Off-query failure
Shared-state reconstruction
ReGround
Information state reliability
🔎 Similar Papers
2024-05-22Neural Information Processing SystemsCitations: 0