๐ค AI Summary
Existing reinforcement learning (RL)-based feature selection methods for high-dimensional complex data suffer from inefficient subspace exploration and suboptimal downstream task performance due to the โone-featureโone-agentโ paradigm. To address this, we propose HRLFS, a multi-agent hierarchical reinforcement learning framework for feature selection. Its key contributions are: (1) a novel hierarchical agent architecture that replaces conventional flat RL designs; (2) integration of statistical features with LLM-driven semantic representations, coupled with hierarchical clustering to group features by semantic similarity; and (3) interpretable, scalable, progressive collaborative optimization over feature subspaces. Evaluated on multiple benchmark datasets, HRLFS achieves an average 3.2% improvement in classification accuracy, reduces runtime by 41%, and demonstrates strong robustness and cross-domain generalization capability.
๐ Abstract
Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream machine learning task. Among filter, wrapper, and embedded-based approaches, the reinforcement learning (RL)-based subspace exploration strategy provides a novel objective optimization-directed perspective and promising performance. Nevertheless, even with improved performance, current reinforcement learning approaches face challenges similar to conventional methods when dealing with complex datasets. These challenges stem from the inefficient paradigm of using one agent per feature and the inherent complexities present in the datasets. This observation motivates us to investigate and address the above issue and propose a novel approach, namely HRLFS. Our methodology initially employs a Large Language Model (LLM)-based hybrid state extractor to capture each feature's mathematical and semantic characteristics. Based on this information, features are clustered, facilitating the construction of hierarchical agents for each cluster and sub-cluster. Extensive experiments demonstrate the efficiency, scalability, and robustness of our approach. Compared to contemporary or the one-feature-one-agent RL-based approaches, HRLFS improves the downstream ML performance with iterative feature subspace exploration while accelerating total run time by reducing the number of agents involved.