🤖 AI Summary
This study addresses the limitations of historical reputation in multi-agent debate, where it poorly predicts behavior on novel tasks and remains vulnerable to adversarial attacks. To overcome these challenges, this work proposes MiniRep, a system that innovatively aggregates real-time task behavior with long-term reputation dynamically. Furthermore, it introduces a taxonomy-based defense mechanism to prevent homogeneous response groups from dominating decision-making, thereby mitigating compound threats involving reputation manipulation and software mutation. Evaluated in a 10-agent heterogeneous setting on the MATH benchmark, MiniRep significantly outperforms conventional baseline methods across all 28 attack configurations. These results demonstrate that the proposed approach effectively enhances both the robustness and security of multi-agent collaboration under complex adversarial conditions.
📝 Abstract
Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration.
We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by this taxonomy, MiniRep evaluates agents based on both their behavior on the current task and their reputation over time, while preventing groups of agents with highly similar responses from dominating the final decision. We assess MiniRep across diverse tasks, LLM-agent compositions, corruption placements, and attack types drawn from our taxonomy. Our experimental results show that, MiniRep outperforms both conventional MAD aggregation and conventional reputation-based approaches on MATH no matter being attacked or not. Also, under a heterogeneous 10-agent setting on MATH, MiniRep outperforms all baselines in all 28 attack conditions.