🤖 AI Summary
This study addresses the failure of conformal prediction caused by the non-exchangeability of candidate distributions and the prohibitive evaluation costs in reinforcement learning (RL)-driven hardware-aware neural architecture search (NAS). To overcome these challenges, we propose an online Adaptive Conformal Inference (ACI) mechanism. By replacing static, one-shot quantile estimation with a tuning-free, locally adaptive online feedback control loop, the proposed method effectively mitigates violations of calibration assumptions induced by dynamic RL policy shifts, thereby ensuring robust coverage guarantees during architecture pruning. Experiments across three architecture families demonstrate that ACI tracks target coverage with precision on the order of 10⁻³ while reducing evaluation overhead by 25%–50% without compromising accuracy. The proposed approach significantly outperforms both static calibration and Gaussian process baselines.
📝 Abstract
Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction $δ$ are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the surrounding reinforcement-learning (RL) loop violates: the policy's proposals improve as search proceeds and, in layer-by-layer construction, shift within every episode. We replace one-shot quantile estimation with online feedback control (Adaptive Conformal Inference, with tuning-free, locally-adaptive, and group-conditional variants), restoring steerable coverage: dialing the target delivers it, monotonically and reproducibly, for arbitrary sequences. Across three neural-network architecture families and both single-step and sequential search (three seeds), it tracks every requested level to within ${\sim}10^{-3}$ while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request. Finally, used as an acquisition function on one constrained testbed, the same optimistic bound beats random search, a gain that fixed optimism already carries and online calibration sharpens. The source code is available at https://github.com/Vicomtech/rl-hw-nas.