🤖 AI Summary
This study addresses the absence of a unified framework for reliably evaluating performance drivers in large language model–based financial multi-agent systems. It proposes a four-dimensional taxonomy to formulate the falsifiable Coordination Priority Hypothesis (CPH) and introduces a novel metric, the Coordination Break-Even Point under Net Trading Costs (CBS). Through a systematic literature review, structured categorization, and bias analysis—supplemented by empirical comparisons across twelve multi-agent systems and two single-agent baselines—the work identifies five categories of evaluation flaws that induce sign reversals in reported returns. Based on these findings, the study establishes minimum evaluation standards to advance methodological rigor and promote a standardized assessment paradigm, thereby laying the groundwork for trustworthy and reproducible research in financial multi-agent systems.
📝 Abstract
Multi-agent systems based on large language models (LLMs) for financial trading have grown rapidly since 2023, yet the field lacks a shared framework for understanding what drives performance or for evaluating claims credibly. This survey makes three contributions. First, we introduce a four-dimensional taxonomy, covering architecture pattern, coordination mechanism, memory architecture, and tool integration; applied to 12 multi-agent systems and two single-agent baselines. Second, we formulate the Coordination Primacy Hypothesis (CPH): inter-agent coordination protocol design is a primary driver of trading decision quality, often exerting greater influence than model scaling. CPH is presented as a falsifiable research hypothesis supported by tiered structural evidence rather than as an empirically validated conclusion; its definitive validation requires evaluation infrastructure that does not yet exist in the field. Third, we document five pervasive evaluation failures (look-ahead bias, survivorship bias, backtesting overfitting, transaction cost neglect, and regime-shift blindness) and show that these can reverse the sign of reported returns. Building on the CPH and the evaluation critique, we introduce the Coordination Breakeven Spread (CBS), a metric for determining whether multi-agent coordination adds genuine value net of transaction costs, and propose minimum evaluation standards as prerequisites for validating the CPH.