Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of execution inconsistency despite correct shared memory in multi-agent systems. By formalizing the concept of "execution consistency," it reveals that retained records alone are insufficient to determine task compliance. To this end, we propose CAVERT, a framework that introduces explicit evidence conditioning to disambiguate opposing label pairs and leverages natural language annotations to extract relational evidence from logs, enabling precise diagnosis and recovery of execution states. Evaluated across twelve benchmarks, CAVERT achieves superior diagnostic performance compared to existing methods. Furthermore, in four experimental environments, its recovery capability consistently outperforms both large language models and rule-based baselines, effectively identifying and repairing execution violations.
📝 Abstract
Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge task duties. We define execution consistency through duties governing state use, information handoffs, and final-state agreement, with explicit evidence conditions for judging fulfillment. Our core claim is that identical retained records can correspond to compliant and violating executions under the same task rule. Controlled removal of evidence such as receipt, action dependence, or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoration separates 97.9% of the merged pairs. Natural-log annotations identify the defined violations in actual executions. However, existing logs do not always explicitly represent the execution relationships needed for these judgments. To assess the definition's practical value, we use CAVERT, a framework for consistency diagnosis and recovery, to extract supported relationships from logs and apply these criteria. It consistently outperforms contract-prompted LLM and rule-based baselines in diagnosis across all 12 benchmark-executor settings. Under the same gate and executor limits, it also outperforms rule-guided recovery in all four evaluated environments. These findings identify execution evidence that agent-memory and execution interfaces should preserve for reliable judgment.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
Execution Consistency
Shared Memory
Failure Diagnosis
Evidence Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Execution Consistency
Multi-Agent Systems
CAVERT Framework
Consistency Diagnosis
Evidence Conditions
🔎 Similar Papers
No similar papers found.