Score
Design and implement feature-selection systems that use reinforcement learning policies (often trained with PPO) to sequentially choose which features to acquire or keep for downstream models; this includes specifying state and action representations, reward functions that trade off prediction accuracy and acquisition or computational cost, training the RL agent, and evaluating the informativeness and efficiency of selected feature subsets.
To address the low efficiency, poor generalization, and weak interpretability of feature selection in high-dimensional machine learning, this paper proposes an automated feature engineering framework based on reinforcement learning (RL). We introduce Monte Carlo Reinforcement Feature Selection (MCRFS), a novel RL algorithm instantiated in three architectures: single-agent, dual-agent cooperative, and cascaded multi-stage RL. A hybrid state representation is designed, combining sequential scanning with convolutional autoencoders; additionally, we integrate bandit-based action selection, early stopping, and hierarchical reward shaping. Extensive experiments across diverse high-dimensional benchmark datasets demonstrate that our method significantly improves feature selection quality and computational efficiency, while enhancing downstream model generalization and interpretability. Empirical results show consistent superiority over conventional feature engineering approaches, including filter, wrapper, and embedded methods.
This work addresses the trade-off between predictive performance and feature acquisition cost by proposing a dynamic, sample-level feature selection method. It formulates active feature selection as a Markov decision process with variable state dimensions and employs reinforcement learning to sequentially recommend the most cost-effective next feature. To enhance scalability and policy simplicity, the approach innovatively integrates a heuristic feature combination exploration strategy tailored for large-scale data and a post-fitting regularization mechanism, effectively compressing decision paths. Experimental results on four binary classification datasets—ranging up to 56 features and 4,500 samples—demonstrate that the proposed method achieves higher accuracy while significantly outperforming existing approaches in terms of decision efficiency and model compactness.
Traditional static feature exclusion strategies fail to mitigate model bias arising from hidden dependencies. This paper proposes an end-to-end reinforcement learning framework that jointly optimizes fairness-aware automated feature selection and predictive modeling. Methodologically, we formulate a multi-objective reward function balancing accuracy and fairness; design a policy-gradient-based action space over feature subsets; and incorporate dynamic regularization and ensemble fusion for adaptive feature selection. Experiments across multiple benchmark datasets demonstrate that our approach maintains high predictive accuracy while significantly reducing bias—e.g., Equalized Odds difference decreases by up to 42%—and improves model generalization and robustness. The core contribution lies in the first integration of fairness constraints directly into the feature selection decision process, thereby overcoming the limitations of static exclusion.
Existing reinforcement learning (RL)-based feature selection methods for high-dimensional complex data suffer from inefficient subspace exploration and suboptimal downstream task performance due to the “one-feature–one-agent” paradigm. To address this, we propose HRLFS, a multi-agent hierarchical reinforcement learning framework for feature selection. Its key contributions are: (1) a novel hierarchical agent architecture that replaces conventional flat RL designs; (2) integration of statistical features with LLM-driven semantic representations, coupled with hierarchical clustering to group features by semantic similarity; and (3) interpretable, scalable, progressive collaborative optimization over feature subspaces. Evaluated on multiple benchmark datasets, HRLFS achieves an average 3.2% improvement in classification accuracy, reduces runtime by 41%, and demonstrates strong robustness and cross-domain generalization capability.
This paper addresses the distribution shift induced by active feature acquisition (AFA) deployment and formally introduces the “Active Feature Acquisition Performance Evaluation” (AFAPE) task—the first of its kind. Under the assumptions of no direct effect (NDE) and no unobserved confounding (NUC), we propose a semi-offline reinforcement learning framework that relaxes the conventional strong positivity assumption. Within this framework, we develop three novel estimators grounded in missing-data mechanisms: Direct Method (DM), Inverse Probability Weighting (IPW), and Doubly Robust Learning (DRL). We establish their consistency and asymptotic normality theoretically. Empirical evaluation demonstrates that our methods significantly improve estimation accuracy and robustness under time-varying feature availability, outperforming existing baselines across diverse benchmarks.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
This work addresses the high computational complexity and inefficiency of value function approximation in high-dimensional structured Markov decision processes (MDPs). By revealing the low-dimensional geometric structure of decision tessellations induced by optimal policies, the authors propose a boundary-driven policy approximation method that directly learns policy regions rather than value functions. They further introduce a policy loss decomposition mechanism that quantitatively links performance degradation to action boundary errors. Evaluated on inventory control and queue admission tasks, the proposed approach significantly reduces policy error and value gap compared to standard reinforcement learning baselines, achieving faster error convergence and enhanced training stability.
This study addresses the limited robustness of deep reinforcement learning agents in real-time strategy games when facing out-of-distribution opponents by proposing a hierarchical architecture that decouples strategic command generation from unit-level control. At the lower level, a constrained instruction-conditioned executor handles unit control, complemented by a policy estimation mechanism grounded in in-game observations rather than opponent identification. At the upper level, a Thompson sampling-based multi-armed bandit serves as a strategist to dynamically select discrete commands. Experimental evaluations in the MicroRTS environment demonstrate that this system achieves significantly higher win rates against most strong adversaries compared to a flat Proximal Policy Optimization baseline, effectively enhancing cross-opponent generalization capabilities.
本文探讨了如何利用强化学习解决运筹学中的动态决策问题,通过与传统方法结合来提高解决方案的质量和效率,并为未来研究提供了路线图。
This study addresses the challenges of overfitting, high computational cost, and limited scalability of traditional methods in feature selection for high-dimensional bioinformatics data. To this end, we propose a large language model (LLM)-in-the-loop reinforcement learning framework that formulates feature selection as a sequential decision-making task. Specifically, the framework leverages LLMs to guide the exploration of the search space and introduces a hybrid reward mechanism integrating domain knowledge with data-driven evaluation to optimize policies. Furthermore, natural language generation techniques are incorporated to enhance interpretability. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods across diverse datasets, achieving consistent improvements in downstream model performance alongside rapid convergence.