From language-model stock rankings to testable economic rules: A computational audit

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unresolved stability, reproducibility, and testability of economic regularities when applying large language models (LLMs) to China A-share stock ranking. It proposes a rigorous evaluation framework that freezes development-period preferences and transfers them to unseen months, employing multi-model comparisons and linear rule fitting alongside Spearman correlation analysis, HAC adjustments, bootstrap testing, and portfolio backtesting for robustness auditing under multiple-hypothesis correction. The findings reveal that compact linear rules can closely approximate aggregated rankings with correlations exceeding 0.9, yet exhibit weak predictive power at the individual stock level. Furthermore, after statistical corrections, LLM-generated rankings fail to produce significant excess returns. This work establishes a rigorous empirical benchmark for AI-driven quantitative investment.
📝 Abstract
We test the stability, reproducibility and investment outcomes of language-model stock rankings. Four models and five numerical comparators share a portfolio engine over 72 monthly holding periods in the Shanghai Stock Exchange (SSE) 50, China Securities Index (CSI) 300 and CSI 500. Rankings use nine characteristics, and five repeated SSE 50 runs measure variation under identical inputs. Linear rules fitted to development-period model preferences are frozen before unseen-month, larger-pool and controlled-intervention tests. Their mean Spearman agreement with model rankings is 0.923-0.984 in the SSE 50 and 0.795-0.985 after transfer. Aggregate rank-change error falls relative to a zero-change prediction in twelve archived feature-group comparisons and eight matched single-feature comparisons, with Holm adjustments applied in separate nine-plus-three and six-plus-two families. Prediction of individual entries and exits remains weak (event Jaccard 0.000-0.125). Historical mean model compound annual growth rates range from 6.26% to 11.17%. At 10 basis points per side and six-month blocks, the twelve-comparison model-minus-rule return family and factor-controlled associations yield no adjusted finding. Three higher-cost, twelve-month-block comparisons favor a Terra rule within their twelve-test slices, with no adjusted finding across the full 144-test sensitivity grid. Two input-intervention batches totaling 5,184 responses supply paired intervention-return tests. The three-model batch has no bootstrap-adjusted finding at the primary block length; a Luna row-order effect appears under heteroskedasticity- and autocorrelation-consistent (HAC) adjustment within its three-test family but not in a pooled 24-test adjustment. Compact rules approximate aggregate rankings; return conclusions depend on comparison families, uncertainty methods and tie priorities.
Problem

Research questions and friction points this paper is trying to address.

language-model stock rankings
reproducibility
investment outcomes
economic rules
computational audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computational Audit
Language Model Stock Ranking
Linear Rule Extraction
Portfolio Engine
Multiple Hypothesis Testing
🔎 Similar Papers