Can Language Models Learn to Forecast Stock Prices

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether post-training can enhance the performance of language models in noisy and information-incomplete financial stock price prediction tasks. We construct a time-series stock sandbox environment and train Qwen3-4B by combining supervised fine-tuning (SFT) with proximal policy optimization (PPO). Furthermore, we introduce tool-calling capabilities and terminal reward mechanisms to enable the model to proactively retrieve market evidence for return prediction. Experimental results demonstrate that post-training significantly improves both the model's financial forecasting capability and its market exploration strategies, such as increasing ranking queries. The resulting AURA-4B model doubles its directional magnitude score to 43.31 and achieves a conditional magnitude consistency of 66.2%, yielding performance comparable to frontier models.
📝 Abstract
Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
Problem

Research questions and friction points this paper is trying to address.

stock price forecasting
language models
financial prediction
post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training
Proximal Policy Optimization (PPO)
Supervised Fine-Tuning (SFT)
Financial Forecasting
Tool Use
🔎 Similar Papers
No similar papers found.