Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning

πŸ“… 2025-05-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Traditional feature selection methods for high-dimensional genomic data neglect biological pathway structures, leading to unstable and biologically uninterpretable results. To address this, we propose a two-stage collaborative reinforcement learning framework that integrates statistical rigor with pathway prior knowledge. Our approach innovatively couples multi-agent reinforcement learning (MARL) with KEGG pathway annotations, employs graph neural networks (GNNs) to model gene–gene interactions, and designs a composite reward function balancing predictive accuracy and pathway coverage. We further introduce shared memory and a centralized critic to enable coordinated agent optimization. Evaluated across multiple gene expression datasets, our method achieves an average 4.2% improvement in disease classification accuracy, enhances pathway enrichment significance by 3.8Γ—, and significantly improves cross-dataset stability and biological consistency.

Technology Category

Machine Learning: Graph-based Machine LearningMultiagent Systems: Multiagent LearningSearch and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
πŸ“ Abstract
Gene selection in high-dimensional genomic data is essential for understanding disease mechanisms and improving therapeutic outcomes. Traditional feature selection methods effectively identify predictive genes but often ignore complex biological pathways and regulatory networks, leading to unstable and biologically irrelevant signatures. Prior approaches, such as Lasso-based methods and statistical filtering, either focus solely on individual gene-outcome associations or fail to capture pathway-level interactions, presenting a key challenge: how to integrate biological pathway knowledge while maintaining statistical rigor in gene selection? To address this gap, we propose a novel two-stage framework that integrates statistical selection with biological pathway knowledge using multi-agent reinforcement learning (MARL). First, we introduce a pathway-guided pre-filtering strategy that leverages multiple statistical methods alongside KEGG pathway information for initial dimensionality reduction. Next, for refined selection, we model genes as collaborative agents in a MARL framework, where each agent optimizes both predictive power and biological relevance. Our framework incorporates pathway knowledge through Graph Neural Network-based state representations, a reward mechanism combining prediction performance with gene centrality and pathway coverage, and collaborative learning strategies using shared memory and a centralized critic component. Extensive experiments on multiple gene expression datasets demonstrate that our approach significantly improves both prediction accuracy and biological interpretability compared to traditional methods.
Problem

Research questions and friction points this paper is trying to address.

Gene selection in high-dimensional genomic data for disease understanding
Integrating biological pathway knowledge with statistical gene selection
Improving gene selection accuracy and biological interpretability via reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pathway-guided pre-filtering with KEGG and statistics
Multi-agent reinforcement learning for gene selection
GNN-based state and reward for biological relevance
πŸ”Ž Similar Papers
No similar papers found.