AlphaOpsBench: Benchmarking End-to-End Alpha Strategy Operationalization in Prediction Markets

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that large language models face in achieving end-to-end translation from abstract concepts to executable code when generating prediction market strategies, primarily due to unrealistic asset standardization assumptions. To this end, we construct a million-scale Polymarket data benchmark and propose a source-preserving strategy logging and lifecycle tracking framework that rigorously decouples strategy fidelity, behavioral validity, and historical executability, systematically comparing direct generation with staged design paradigms. Our findings reveal that strict end-to-end effectiveness remains exceedingly low and that replaying strategies is substantially easier than executing them faithfully. Furthermore, we demonstrate that existing models lack economic decision-making stability in real-world strategy generation. Ultimately, this work provides a critical evaluation benchmark and methodological insights for the automated generation of quantitative trading strategies.
📝 Abstract
Large language models increasingly generate quantitative trading strategies, yet existing benchmarks assume standardized assets, numerical features, or directly compilable strategy representations---assumptions that prediction-market strategies violate, since a coarse idea may leave the traded outcome, causal information source, signal definition, threshold, sizing, order policy, exit, and settlement behavior unspecified. We introduce \textsc{AlphaOpsBench}, which evaluates end-to-end operationalization from source-grounded economic hypotheses to auditable executable programs over 581 source-preserving strategy records and a lifecycle-scale Polymarket dataset with 1.28 million binary markets, 183.6 million cleaned executions, settlement evidence, and limit-order-book history, comparing Direct generation against a Staged design-then-code protocol. In a corrected independent-generation study over 36 controlled tasks and 24 preregistered real strategies, strict end-to-end validity remains rare: Direct and Staged obtain 35/180 and 20/180 canonical passes on the controlled cohort and no confirmed pass on the real cohort, and repeated generations vary substantially in model-owned economic choices. By contrast, 775,725 of 783,655 scheduled historical replays complete, showing that replayability is a far weaker property than source-faithful operationalization. Financial outcomes depend on the declared execution model and available historical evidence, and fee and liquidity experiments show that execution costs alter subsequent trading paths rather than acting only as ex-post deductions. \textsc{AlphaOpsBench} thus separates strategy fidelity, behavioral validity, historical executability, and financial performance in an evidence-aware benchmark for LLM-based quantitative research in prediction markets.
Problem

Research questions and friction points this paper is trying to address.

prediction markets
quantitative trading strategies
large language models
end-to-end operationalization
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prediction Markets
Large Language Models
Strategy Operationalization
Benchmark
Quantitative Trading
🔎 Similar Papers
H
Huaiyu Jia
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
M
Mingxuan Zhao
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
J
Jincheng Gao
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Zifan Peng
Zifan Peng
Ph.D. Candidate at HKUST(GZ)
DeFiTrustworthy AI
Wentao Zhang
Wentao Zhang
Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsctime-resolved
S
Siguang Li
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Shuo Sun
Shuo Sun
Johns Hopkins University