🤖 AI Summary
This study addresses the persistent reasoning errors in multi-agent systems caused by memory contamination and scope collapse. To mitigate these issues, we propose a dynamic reliability governance framework that transitions from static storage paradigms to a dynamic governance loop. Specifically, an asymmetric consolidation mechanism is introduced to preserve coordination scope, complemented by multi-scale confidence-aware evaluation with paired replay verification. Furthermore, downstream impact optimization is leveraged to refine high-risk knowledge auditing strategies under limited budgets. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across most benchmarks, outperforming the strongest baseline by 10.23 percentage points. Ablation studies confirm the critical contribution of the scope protection mechanism, whose removal incurs a substantial accuracy degradation of 16.89 percentage points.
📝 Abstract
Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly supported, or spuriously successful procedures become recurring components of future reasoning. Multi-agent execution introduces an additional structural risk. Scope collapse occurs when procedural knowledge escapes the coordination scope in which it was shown effective and is repeatedly reused at incompatible decision levels, allowing local errors to influence cascades of downstream decisions. Meanwhile, task-level failures provide ambiguous supervision because they rarely reveal which recalled knowledge was responsible. We introduce AECG, a framework for asymmetric experience consolidation and governance for multi-agent systems. AECG turns memory from static experience storage into a dynamic reliability-governance loop, preserving coordination scope and using multi-scale, confidence-aware reliability to detect degradation. It then combines degradation with downstream impact to prioritize high-risk knowledge under a bounded review budget, applies targeted interventions, and reactivates revised skills only after paired replay. Across three multi-agent frameworks and four benchmarks, AECG achieves the best score in 11 of 12 framework--benchmark settings and improves over the strongest competing memory method by as much as 10.23 percentage points; removing scope preservation reduces accuracy by up to 16.89 points. AECG thereby reframes multi-agent memory from passive accumulation into auditable reliability governance. Code is available at https://github.com/fenhg297/AECG