arXiv · 2610.03127
Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS
Abstract
Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction $δ$ are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the surrounding reinforcement-learning (RL) loop violates: the policy's proposals improve as search proceeds and, in layer-by-layer construction, shift within every episode. We replace one-shot quantile estimation with online feedback control (Adaptive Conformal Inference, with tuning-free, locally-adaptive, and group-conditional variants), restoring steerable coverage: dialing the target delivers it, monotonically and reproducibly, for arbitrary sequences. Across three neural-network architecture families and both single-step and sequential search (three seeds), it tracks every requested level to within ${\sim}10^{-3}$ while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request. Finally, used as an acquisition function on one constrained testbed, the same optimistic bound beats random search, a gain that fixed optimism already carries and online calibration sharpens. The source code is available at https://github.com/Vicomtech/rl-hw-nas.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pedro Brandimarte, Nerea Aranjuelo, Marcos Nieto, Oihana Otaegui. 2026-10-02. Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS. https://arxiv.org/abs/2610.03127
Cite the original work for its findings. Save a collection to share your selection of sources.