🤖 AI Summary
This paper frames asset pricing as a multiclass classification task, predicting whether individual stocks will outperform or underperform the market. Method: Supervised learning models generate out-of-sample trading portfolios, and a novel sequential binomial test rigorously evaluates over 3.34 million predictions for statistical significance—constituting the first systematic validation that historical information contains non-random, machine-learnable predictability. Contribution/Results: (1) Multiple classifiers pass stringent statistical tests, confirming short-term market inefficiency; (2) Model uncertainty—quantified by prediction probabilities—significantly affects portfolio performance, with high-confidence predictions yielding superior economic returns; (3) The constructed portfolios consistently outperform benchmarks out-of-sample. Collectively, this work establishes a reproducible, statistically verifiable methodological framework and empirical foundation for machine learning–driven quantitative investing.
📝 Abstract
We design a novel framework to examine market efficiency through out-of-sample (OOS) predictability. We frame the asset pricing problem as a machine learning classification problem and construct classification models to predict return states. The prediction-based portfolios beat the market with significant OOS economic gains. We measure prediction accuracies directly. For each model, we introduce a novel application of binomial test to test the accuracy of 3.34 million return state predictions. The tests show that our models can extract useful contents from historical information to predict future return states. We provide unique economic insights about OOS predictability and machine learning models.