Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

📅 2026-06-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current LLM-driven trading research lacks standardized execution assumptions and reproducibility criteria, hindering cross-study comparisons and economic interpretability. This work systematically reviews 30 related studies and introduces the first evidence matrix encompassing execution semantics, turnover handling, and temporal control to evaluate transparency across dimensions such as data recency, backtest partitioning, and transaction cost modeling. Through bibliometric coding, methodological analysis of backtesting practices, and friction sensitivity experiments on ten stocks, the study quantifies how execution details compress strategy returns. It reveals that most papers inadequately disclose execution assumptions and proposes a standardized reporting framework emphasizing execution transparency as critical for result credibility, advocating for stricter community-wide standards of realism and reproducibility.
📝 Abstract
Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling. This article presents a targeted topical review and reproducibility audit of execution realism in LLM-based trading research. A coded evidence matrix covering 30 trade-relevant primary studies is used to assess point-in-time controls, split transparency, held-out evaluation, cost and turnover treatment, execution semantics, universe definition, and artifact release. Across the audited sample, architecture reporting is generally clearer than the evaluation assumptions needed to judge whether a trading result is economically interpretable or reproducible. A 10-equity worked example is included only as a methodological scaffold to illustrate how explicit friction and timing choices can materially compress active-strategy results. The main conclusion is that the next useful step for LLM trading research is not only better agent design, but also clearer reporting standards for execution realism, reproducibility, and evaluation comparability.
Problem

Research questions and friction points this paper is trying to address.

LLM-based trading
execution realism
reproducibility
evaluation comparability
transaction cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

execution realism
reproducibility
LLM-based trading
evaluation comparability
transaction cost modeling
🔎 Similar Papers
No similar papers found.