🤖 AI Summary
This study addresses the high computational cost of fixed-depth inference in pretrained tabular foundation models and their susceptibility to context contamination, which degrades anomaly detection performance. We propose a depth-adaptive early exit mechanism, revealing for the first time its critical role in enhancing detection accuracy beyond mere inference acceleration. By designing a plug-and-play router based on privileged information to dynamically select optimal exit layers, our approach effectively mitigates context contamination. Experiments demonstrate that this method improves average detection performance by 4.7–7.3%, with gains reaching 13–21% under severe contamination scenarios. Furthermore, it achieves up to 1.8× inference speedup while recovering 45% of the potential performance benefits.
📝 Abstract
Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we uncover a surprising benefit: exiting at the optimal intermediate layer can also improve detection performance on diverse real-world benchmarks by 4.7-7.3% on average, consistent across three distinct foundation models. First, we investigate the factors driving these gains, and identify a key mechanism: context pollution, i.e., the presence of outliers among in-context samples. Our analysis reveals that nearby in-context samples exert increasing influence on query predictions at greater depths, consistent with a retrieval-based view of these models. In effect, early-exit alleviates the adverse effects of retrieving accurate-yet-polluted neighbors, with gains of 13-21% when context pollution matches the natural outlier rate. Motivated by these findings, we pretrain a plug-in router to select a dataset-specific exit layer, using query outlier labels as privileged information available only during router training. The router operates post hoc, leaving the base model parameters and prediction head unchanged. Experiments on three large real-world benchmarks show that, on clean context, the router recovers up to 45% of the oracle gain with up to 1.8x speedup across three pretrained backbones, with larger gains as context pollution increases.