Active Feature Acquisition With Incomplete Training Data

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the significant performance degradation of multi-step policies in active feature acquisition caused by missing training data. By investigating active feature acquisition under incomplete data conditions, the authors employ statistical learning theory to characterize the decay of multi-step acquisition value under the Missing Completely At Random (MCAR) assumption. Furthermore, this work systematically analyzes the theoretical implications of three mitigation strategies: aliasing, filtering, and generative recovery. It is demonstrated that generative recovery effectively compensates for performance loss, substantially outperforming the filtering approach whose complexity grows exponentially. A benchmark framework is constructed to validate that missing data primarily impairs multi-step methods and that generative recovery significantly enhances their performance. The source code has been made publicly available.
πŸ“ Abstract
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by formalizing the trade-off between feature cost and predictive performance during sequential feature selection. However, prior AFA work largely assumes access to complete training data, an assumption that is often violated in practice. Here we study AFA with Incomplete Training Data (AFA-ITD), showing that under missing completely at random (MCAR) data, one-step acquisition values remain unchanged, whereas multi-step values can decrease. We analyze three approaches to learning from incomplete data: aliasing, filtering, and generative restoration. We show that filtering can require a number of training instances scaling exponentially with the dimension, whereas generative restoration scales exponentially with the acquisition budget. We empirically test our theory on a controlled experiment and across common AFA datasets and find that missingness mainly damages methods that exploit multi-step acquisitions and that generative restoration is able to recover lost performance in many experiments. Code is available at https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data.
Problem

Research questions and friction points this paper is trying to address.

Active Feature Acquisition
Incomplete Training Data
Missing Data
Sequential Feature Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Feature Acquisition
Incomplete Training Data
Generative Restoration
Missing Completely at Random
Sequential Feature Selection
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
R
Reza Rezvan
Department of Computer Science and Engineering, Chalmers University of Technology & University of Gothenburg
V
Valter SchΓΌtz
Department of Computer Science and Engineering, Chalmers University of Technology & University of Gothenburg
H
Han Wu
Department of Computer Science and Engineering, Chalmers University of Technology & University of Gothenburg
Linus Aronsson
Linus Aronsson
Chalmers University of Technology
Machine Learning
Morteza Haghir Chehreghani
Morteza Haghir Chehreghani
Chalmers University of Technology
Artificial IntelligenceMachine LearningData ScienceDeep Learning