LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that evaluating large language model (LLM) agents solely on outcomes conflates capability with performance, hindering the diagnosis of mechanistic utilization differences in dynamic environments. To overcome this, we leverage real-time financial markets to construct a process-aware benchmark, transforming markets from leaderboards into diagnostic environments. Through continuous trajectory testing, exhaustive decision tracking, and multidimensional mechanism configuration comparisons, we enable closed-loop interaction between agents and financial data. Our findings reveal a significant divergence between outcomes and capabilities, demonstrating that mechanism access, effective utilization, and downstream performance are not interchangeable. Furthermore, we quantify capability gaps across five frontier models, precisely identifying specific bottlenecks in dimensions such as tool calling and memory, thereby achieving fine-grained diagnostics of agent capabilities.
📝 Abstract
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability
Problem

Research questions and friction points this paper is trying to address.

LLM agent evaluation
process-aware benchmark
outcome-capability gap
evolving environments
capability diagnostics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process-Aware Evaluation
LLM Agents
Live Financial Markets
Decision Traces
Capability Diagnostics
🔎 Similar Papers
No similar papers found.
J
Jun Zhao
National University of Singapore
L
Leiming Fu
Fudan University
Y
Yanbo Wen
Fudan University
Y
Yiding Wang
Fudan University
Xuantong Liu
Xuantong Liu
Fudan University
Y
Yang Shu
Fudan University
Y
Yuyang Lu
Fudan University
X
Xuanran Xing
Fudan University
J
Jingqi Tong
Fudan University
Hao Xu
Hao Xu
The University of Sydney; GAI
Qi Zhang
Qi Zhang
Fudan University
SAGINsatellite routing
X
Xuanjing Huang
Fudan University