🤖 AI Summary
This study addresses the challenge of strategy selection and portfolio construction arising from the vast number of available option contracts by proposing a large language model (LLM)-based agent framework. Built upon Qwen3.8-27B, the framework integrates supervised fine-tuning with reinforcement learning to abstract numerous options into a strategy-level decision space, overcoming fixed structural constraints while executing dynamic portfolio switching through deterministic parsing algorithms. The research further reveals an asymmetric effect of news data during model training. Six-month out-of-sample backtesting demonstrates that the system achieves a total return of 18.3%, a Sharpe ratio of 1.60, and a maximum drawdown of only 8.96%, thereby validating the effectiveness of LLM-driven option strategy generation.
📝 Abstract
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.