MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing multi-agent benchmarks focus solely on final outcomes, obscuring the mechanisms underlying collaborative gains and degradation. This study reveals that while interaction improves weak proposals, it frequently degrades strong ones. To address this, we propose MASTraceBench, a multi-level metric framework that traces proposal trajectories to diagnose collaborative processes. Furthermore, we introduce CLEARS, a method that replaces holistic exchanges with claim-level evaluation to effectively prevent the degradation of strong initial proposals. Experimental results demonstrate that CLEARS achieves the highest collaborative gains on five out of six tasks, significantly enhancing both the preservation and refinement of strong proposals.
📝 Abstract
LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
Collaboration Gain
Proposal Trajectories
Benchmark
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Systems
Proposal Trajectories
Collaboration Gain
Benchmark
Claim-level Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yapeng Li
Harbin Institute of Technology, Harbin, China
S
Songze Li
Harbin Institute of Technology, Harbin, China
Shuang Yu
Shuang Yu
Harbin Institute of Technology, Harbin, China
Jing Yu
Jing Yu
Northwestern University
SustainabilityLife Cycle AnalysisTransportation ManagementOperations Research
Z
Zhixin Liu
Harbin Institute of Technology, Harbin, China
L
Liqiang Wen
Peking University, Beijing, China
Tonghua Su
Tonghua Su
Professor of Harbin Institute of Technology
pattern recognitioncharacter recognitionmachine learningsoftware engineering