The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the inference costs of large language models (LLMs) can be translated into net return improvements in quantitative trading. Treating LLM inference as an economic intervention for the first time, we systematically evaluate the post-cost economic value of varying reasoning intensities in U.S. equity markets. By controlling variables and comparing models such as DeepSeek within a framework integrating multi-source data inputs, portfolio construction, and backtesting, we assess actual trading performance. Our empirical findings reveal a non-monotonic and unstable relationship between reasoning investment and trading outcomes. Specifically, additional inferential computation does not reliably enhance net returns but instead exacerbates decision volatility. This work provides critical evidence delineating cost-benefit boundaries for AI-driven financial investment.
📝 Abstract
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Test-Time Reasoning
Portfolio Returns
Stock Trading
Economic Value
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Reasoning
LLM Trading
Portfolio Returns
Controlled Study
Nonmonotonic Performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiayi Chen
Department of Computer Science, New Jersey Institute of Technology
Guiling Wang
Guiling Wang
University of Connecticut
Water CycleClimate ChangeClimate ExtremesEcosystemLand-Atmosphere Interactions