🤖 AI Summary
This study addresses the prediction performance bottleneck caused by missing auxiliary information during deployment and limited annotation budgets. To overcome this, it proposes ALCATRAs, a unified framework that integrates data acquisition, surrogate construction, and downstream prediction. Through task selection strategies and surrogate learning, the framework adaptively allocates resources under cost constraints to acquire critical auxiliary information, leveraging active learning, multi-task learning, and surrogate model transfer for efficient optimization. Theoretically, the authors demonstrate that ALCATRAs effectively reduces the prediction error bound. Empirical evaluations on benchmarks such as the UCI Heart Disease dataset confirm that the proposed approach significantly improves both sample efficiency and predictive accuracy compared to baseline methods.
📝 Abstract
Many scientific studies allow costly auxiliary information to be collected during data labeling but not at deployment. Examples include diagnostic tests, laboratory assays, and expert evaluations. We study prediction under this deployment asymmetry, where auxiliary variables are selectively acquired during labeling under a budget constraint but systematically unavailable at prediction time, creating a missing-by-design problem that couples data acquisition, surrogate construction, and prediction. In this work, we introduce Active Learning with Cost-Adaptive Task Resource Allocations (ALCATRAs), a unified framework for selectively acquiring auxiliary information under resource constraints and leveraging that information to improve downstream prediction. ALCATRAs consists of two main components: a task-selection policy which strategically selects a sequence of cost-effective tasks for unlabeled data to perform, and a surrogate learning procedure which transfers knowledge from completed tasks to enhance model predictions. In theory, we show the effectiveness of surrogate models and sample-efficient task policies in improving the model's prediction error bound. Simulation studies and an application to the UCI heart disease cohort demonstrate improved sample efficiency of the proposed ALCATRAs framework relative to baselines under the studied settings.