On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of jointly identifying informative features and underlying subspaces in nonstationary, nonlinear regression settings by proposing Entropy-Optimal Manifold Regression (EOMR). EOMR extends entropy-optimal manifold clustering to regression for the first time, integrating information theory with manifold learning to simultaneously optimize key feature selection and the construction of a low-dimensional causal subspace, all while achieving linear computational and memory complexity. Evaluated on high-dimensional chaotic systems such as Lorenz-96 and Hasegawa-Wakatani, EOMR—using only eight parameters—significantly outperforms gradient-boosted trees, deep neural networks, and TabPFN, reducing prediction errors by several orders of magnitude and effectively capturing the dominant dynamics.
📝 Abstract
We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces of relevant features in nonstationary and nonlinear regression problems. It is shown that the proposed extension - that we coin as Entropy-Optimal Manifold Regression (EOMR) - allows a robust learning with linearly-scaling iteration and memory complexities. EOMR is compared to the most complete set of state-of-the-art tools from the Artificial Intelligence (AI) and Machine Learning (ML) that is available to the author, on the very challenging problems from chaotic and fluid dynamics: (i) on predicting the Lorenz-96 systems dynamics in strongly- and very-strongly chaotic regimes (with forcing parameter being $F=8$ and $F=12$, respectively); and, (ii) on a data from the Hasegawa-Wakatani model on the edge of the tokamak plasma. It is demonstrated that the proposed benchmarks (i) and (ii), indeed, are the very challenging problems for the state of the art ML and AI tools - since both the general-purpose gradient boosted random forests and deep neuronal networks, as well as transformer-based AI tools like TabPFN v.03 (more spezialised for large-dimensional small data learning problems) - result in orders of magnitude inferior root mean squared prediction errors, and orders of magnitude larger model complexities, when compared to the EOMR. For a Hasegawa-Wakatani example, EOMR distills a very simple entropy-optimal and skilful description of the leading Essential Orthogonal Function (EOF) dynamics, given by linear, causal and weakly-stationary autoregressive process described by just 8 parameters.
Problem

Research questions and friction points this paper is trying to address.

feature selection
subspace learning
nonlinear regression
nonstationary data
manifold learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy-Optimal Manifold Regression
feature subset selection
nonlinear regression
chaotic dynamics
model complexity reduction
🔎 Similar Papers
No similar papers found.