🤖 AI Summary
This study addresses the paradox in multi-agent LLM systems wherein high mutual trust enhances performance yet exacerbates vulnerability. We present the first operationalization of a hierarchical trust model by constructing a five-layer trust stack. Through generalized-mean composite trust aggregation, regret-minimization-based online weight adaptation, and short-term revocable capability tokens, the framework achieves cross-layer coordination and dynamically trustworthy capability delegation. Theoretically, we prove that this approach breaks the trust–vulnerability paradox. Empirically, experiments confirm the validity of composite trust bounds, yielding a fixed-point error below 0.007, and demonstrate that compromised premise layers trigger automatic revocation within a few interactions.
📝 Abstract
Layered trust models for multi-agent LLM systems remain largely conceptual: they name which dimensions of trust matter but not how layers combine, how their importance is set at runtime, or how trust should govern agent actions. This gap matters because higher inter-agent trust raises task success while also enlarging exposure to exploitation, a tension formalized as the Trust-Vulnerability Paradox. We make a five-layer trust stack operational through three contributions. First, a cross-layer synergy operator propagates prerequisite-layer deficits into de- pendent layers, feeding a generalized-mean composite trust that recovers the weakest-link rule as a limiting case, with provable bounds. Second, per-layer importance weights are grounded in observed failures via a no-regret online estimator that tracks which layer is currently most responsible for harm. Third, Trust- Gated Capability Control issues short-lived, revocable capability grants only when composite trust and the relevant prerequisite layers clear capability-specific thresholds. We prove this mechanism breaks the paradox: a stealthy compromise that inflates behavioral trust while degrading a prerequisite layer cannot escalate privilege, and give a closed-form, provably conservative trust fixed point with a bounded-latency revocation guarantee. A numerical study with a sleeper adversary confirms the predictions: composite-trust bounds hold across 20,000 random draws, the analytic fixed point matches simulation within 0.007, and a compromised prerequisite layer triggers automatic revocation within a few interactions.