From language-model stock rankings to testable economic rules: A computational audit
This study addresses the unresolved stability, reproducibility, and testability of economic regularities when applying large language models (LLMs) to China A-share stock ranking. It proposes a rigorous evaluation framework that freezes development-period preferences and transfers them to unseen months, employing multi-model comparisons and linear rule fitting alongside Spearman correlation analysis, HAC adjustments, bootstrap testing, and portfolio backtesting for robustness auditing under multiple-hypothesis correction. The findings reveal that compact linear rules can closely approximate aggregated rankings with correlations exceeding 0.9, yet exhibit weak predictive power at the individual stock level. Furthermore, after statistical corrections, LLM-generated rankings fail to produce significant excess returns. This work establishes a rigorous empirical benchmark for AI-driven quantitative investment.