Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of conformal prediction caused by the non-exchangeability of candidate distributions and the prohibitive evaluation costs in reinforcement learning (RL)-driven hardware-aware neural architecture search (NAS). To overcome these challenges, we propose an online Adaptive Conformal Inference (ACI) mechanism. By replacing static, one-shot quantile estimation with a tuning-free, locally adaptive online feedback control loop, the proposed method effectively mitigates violations of calibration assumptions induced by dynamic RL policy shifts, thereby ensuring robust coverage guarantees during architecture pruning. Experiments across three architecture families demonstrate that ACI tracks target coverage with precision on the order of 10⁻³ while reducing evaluation overhead by 25%–50% without compromising accuracy. The proposed approach significantly outperforms both static calibration and Gaussian process baselines.
📝 Abstract
Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction $δ$ are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the surrounding reinforcement-learning (RL) loop violates: the policy's proposals improve as search proceeds and, in layer-by-layer construction, shift within every episode. We replace one-shot quantile estimation with online feedback control (Adaptive Conformal Inference, with tuning-free, locally-adaptive, and group-conditional variants), restoring steerable coverage: dialing the target delivers it, monotonically and reproducibly, for arbitrary sequences. Across three neural-network architecture families and both single-step and sequential search (three seeds), it tracks every requested level to within ${\sim}10^{-3}$ while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request. Finally, used as an acquisition function on one constrained testbed, the same optimistic bound beats random search, a gain that fixed optimism already carries and online calibration sharpens. The source code is available at https://github.com/Vicomtech/rl-hw-nas.
Problem

Research questions and friction points this paper is trying to address.

Hardware-aware NAS
Conformal prediction
Reinforcement learning
Exchangeability violation
Coverage control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hardware-aware NAS
Adaptive Conformal Inference
Reinforcement Learning
Online Calibration
Steerable Coverage
🔎 Similar Papers
P
Pedro Brandimarte
Vicomtech Foundation, Basque Research and Technology Alliance (BRTA), Donostia-San Sebastián, Spain
N
Nerea Aranjuelo
Vicomtech Foundation, Basque Research and Technology Alliance (BRTA), Donostia-San Sebastián, Spain
Marcos Nieto
Marcos Nieto
Principal Researcher, Vicomtech
Computer visionDriver AssistanceBayesian inference
Oihana Otaegui
Oihana Otaegui
vicomtech