MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of gray failures, characterized by persistent task quality degradation in decentralized LLM agent networks, by proposing MeshHeal, a self-healing framework. MeshHeal introduces a novel dual-timescale peer review mechanism: a fast timescale enables immediate output correction, while a slow timescale combines adaptive hierarchical review with a relative performance detector to identify persistent degradation and dynamically route agents for isolation or recovery. Furthermore, model-supported evaluation is incorporated to eliminate hidden routing errors. Experimental results demonstrate that this approach achieves higher accuracy (0.839) on benchmarks such as BBH with significantly reduced token consumption (51k versus 115k), effectively supporting both the isolation and lossless recovery of faulty agents.
📝 Abstract
Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.
Problem

Research questions and friction points this paper is trying to address.

gray failures
decentralized multi-agent systems
self-healing
LLM agents
agent degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decentralized LLM Agents
Gray Failures
Self-Healing
Two-Timescale Peer Review
MAS Evaluation