AlphaPADI: Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of formulaic alpha discovery, including the neglect of pool context, inadequate preservation of structural hierarchy, and the non-differentiability of pool-level rewards. To overcome these challenges, we propose AlphaPADI, a novel framework that introduces a pool-aware hierarchical discrete diffusion mechanism. By integrating syntax-constrained initialization, multi-scale structural reconstruction, and preference learning, AlphaPADI transcends the bottleneck of conventional item-wise generation, which fails to exploit complementary information within the pool, thereby enabling the efficient synthesis of complementary alpha pools. Empirical evaluations on Chinese and U.S. stock markets demonstrate that the proposed method significantly outperforms baseline approaches in both predictive performance and portfolio returns, validating the effectiveness of pool-aware generation for financial signal discovery.
📝 Abstract
Formulaic alpha discovery seeks symbolic expressions that predict cross-sectional asset returns. In deployment, multiple formulas are combined into an alpha pool, where each formula is valued through the complementary information it contributes to joint predictive performance. While Reinforcement Learning and Generative Flow Networks have emerged as promising paradigms for generating formulaic alphas, existing frameworks face three related challenges. First, generating formulas individually leaves pool context and inter-formula complementarity outside the generative state. Second, formula-wise generation lacks a unified mechanism for preserving and revising structures at different levels. Third, pool-level rewards jointly reflect predictive performance and redundancy but cannot be differentiated directly through symbolic evaluation to train the generator. To overcome these challenges, we introduce AlphaPADI (Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion), a novel framework built around three components: (1) grammar-constrained buffer initialization that constructs syntactically valid pool candidates, (2) pool-aware hierarchical diffusion that reconstructs complete pools at multiple structural scales under the current pool context, and (3) reward-guided pool refinement that evaluates joint predictive performance and inner diversity, updates the elite buffer, and trains the reverse model through reconstruction and preference learning. Empirical results on the Chinese and U.S. stock markets demonstrate that AlphaPADI outperforms the evaluated baselines in both predictive and portfolio performance, thereby validating pool-aware generation as an effective framework for automated alpha discovery.
Problem

Research questions and friction points this paper is trying to address.

Formulaic Alpha Discovery
Alpha Pool
Inter-formula Complementarity
Hierarchical Generation
Pool-level Rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Formulaic Alpha Discovery
Hierarchical Discrete Diffusion
Pool-Aware Generation
Preference Learning
Quantitative Finance
🔎 Similar Papers
No similar papers found.