🤖 AI Summary
This study addresses the limitation that evaluating large language model (LLM) agents solely on outcomes conflates capability with performance, hindering the diagnosis of mechanistic utilization differences in dynamic environments. To overcome this, we leverage real-time financial markets to construct a process-aware benchmark, transforming markets from leaderboards into diagnostic environments. Through continuous trajectory testing, exhaustive decision tracking, and multidimensional mechanism configuration comparisons, we enable closed-loop interaction between agents and financial data. Our findings reveal a significant divergence between outcomes and capabilities, demonstrating that mechanism access, effective utilization, and downstream performance are not interchangeable. Furthermore, we quantify capability gaps across five frontier models, precisely identifying specific bottlenecks in dimensions such as tool calling and memory, thereby achieving fine-grained diagnostics of agent capabilities.
📝 Abstract
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability