🤖 AI Summary
This study addresses the accuracy-speed trade-off in surrogate modeling by determining precisely when to invoke computationally expensive full-scale simulators. To this end, it proposes a threshold-aware conformal routing framework that integrates conformal prediction, split calibration, and a learned input-dependent scaling mechanism to dynamically adjust prediction interval tightness according to decision boundaries. By introducing a threshold-aware objective function to optimize interval allocation, the method intelligently routes inputs to either the surrogate model or the simulator. The primary contribution lies in significantly reducing simulator invocation rates by 14%–75% compared to standard conformal prediction, effectively mitigating unnecessary computational overhead while preserving distribution-free statistical coverage guarantees.
📝 Abstract
High-fidelity simulations are essential to scientific and engineering design, but can be expensive to run repeatedly. Learned surrogates offer a faster alternative, yet their higher errors may alter downstream decisions. This accuracy-speed tradeoff creates a need to determine whether a surrogate can be used or the full simulator remains necessary. We study decisions determined by whether a scalar quantity of interest lies above or below a fixed threshold. For each input, we use the surrogate when its conformal interval lies entirely on one side of the threshold and route the input to simulation when the interval intersects it. Standard conformal prediction constructs intervals without reference to the downstream decision threshold: even a narrow interval near the threshold can cross it and trigger simulation, whereas a wider interval farther away can remain entirely on one side and require no simulation. We introduce Threshold-Aware Conformal Routing (TACR), which learns an input-dependent scale using a threshold-aware objective that concentrates interval tightness near the decision boundary. Exact split-conformal calibration on held-out data preserves distribution-free marginal coverage, which also upper-bounds the probability of an incorrect threshold decision that is not routed. Across various scientific and engineering datasets, TACR reduces simulator deferrals by 14-75% relative to standard conformal prediction at the same coverage target. Against a variant without threshold-local weighting but with similar predictor accuracy, TACR further reduces deferrals by 10-24% on four datasets. These results show that optimizing interval allocation for routing can reduce simulator calls without weakening the standard conformal guarantee.