🤖 AI Summary
Existing multi-agent benchmarks focus solely on final outcomes, obscuring the mechanisms underlying collaborative gains and degradation. This study reveals that while interaction improves weak proposals, it frequently degrades strong ones. To address this, we propose MASTraceBench, a multi-level metric framework that traces proposal trajectories to diagnose collaborative processes. Furthermore, we introduce CLEARS, a method that replaces holistic exchanges with claim-level evaluation to effectively prevent the degradation of strong initial proposals. Experimental results demonstrate that CLEARS achieves the highest collaborative gains on five out of six tasks, significantly enhancing both the preservation and refinement of strong proposals.
📝 Abstract
LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.