Score
Designs and implements test-time adaptation procedures that compute per-sample co‑occurrence statistics and use them to weight or filter adaptation updates so that samples with atypical or noisy co‑occurrences contribute less to parameter changes. Builds the weighting functions, update rules and diagnostics needed to adapt models without access to source data and analyzes the impact of co‑occurrence weighting on robustness, noise suppression, and adaptation stability.
研究测试时适应(TTA)在CIFAR-10-C数据集上的表现,通过比较BN-Adapt、TENT和EATA三种方法,揭示了在某些条件下TTA可能无效甚至有害。
Test-time adaptation (TTA) suffers from latent model degradation due to distribution shifts, yet lacks effective online monitoring mechanisms. Method: We propose the first online risk monitoring framework for TTA, leveraging confidence sequences to construct a sequential hypothesis test that dynamically estimates model performance using only unlabeled test samples and detects statistically significant performance deterioration in real time. Contribution/Results: By rigorously integrating statistical inference into TTA—enabling unsupervised, online, and falsifiable failure detection—we bridge a critical methodological gap. Extensive experiments across multiple datasets, diverse shift types (e.g., corruption, domain, semantic), and state-of-the-art TTA algorithms demonstrate that our framework triggers alarms with high precision and low latency, substantially enhancing deployment robustness and safety.
Under distribution shift, model performance degradation primarily stems from reliance on non-causal features—those statistically correlated with but not causally related to the target variable. To address this, we propose Causal Pruning, a test-time adaptive framework that explicitly disentangles non-causal components from learned representations. Our method generates causal-invariant perturbations via data augmentation and employs principal component analysis to dynamically identify high-variance, non-causal directions in the representation space. During inference, it applies orthogonal projection to prune features along these directions, thereby suppressing non-causal signals. Crucially, Causal Pruning operates entirely at test time without requiring source-domain labels or domain metadata. Evaluated across multiple real-world out-of-distribution (OOD) benchmarks, it consistently surpasses state-of-the-art methods, demonstrating superior robustness, effective non-causal feature mitigation, and strong generalization under unseen distribution shifts.
To ensure safety in high-risk AI systems, continuous monitoring for abrupt distributional shifts—such as concept drift, covariate shift, and out-of-support shifts—is essential post-deployment. This paper proposes the Weighted Conformal Test Martingale (WCTM), a generalized nonparametric online changepoint detection framework. WCTM is the first method to achieve *anytime-valid* detection of arbitrary distributional shifts while rigorously controlling the false alarm rate; it further supports online adaptation to mild covariate shifts. Its theoretical foundation integrates weighted conformal prediction, anytime-valid inference, and online martingale construction, unifying the modeling of covariate, concept, and support-set shifts. Evaluated on multiple real-world datasets, WCTM significantly outperforms state-of-the-art methods, achieving superior trade-offs between detection sensitivity and false positive rate, and demonstrating distinct responsiveness to both adaptive and non-adaptive shifts.
论文提出使用预测碎片化方法来解决测试时适应性问题,通过测量初始模型与适应后模型之间的不一致性,以减少有害接受区域,提高适应效果。
This work addresses the challenge of verifying covariate balance in covariate shift adaptation by proposing a sequentially valid, anytime-stoppable validation framework. Built upon time-uniform confidence sequences, the method dynamically monitors covariate balance for a pre-specified function class within a prescribed tolerance band and terminates as soon as all target moments fall within this band, thereby certifying balance. Its key contribution lies in providing, for the first time, a locally and absolutely valid certification of balance for any adjustment strategy, supporting data-dependent stopping times while rigorously controlling the probability of erroneous balance confirmation. Integrated with KL-divergence-based drift diagnostics and detection of admissible adjustment regions, experiments demonstrate the method’s superior performance in error rate control, locality with respect to function classes, effectiveness in drift diagnosis, and guaranteed coverage in conformal prediction.
This work addresses the challenge of distribution shift between training and test phases by proposing the first online test-time adaptation framework based on state space modeling. The method learns an initial model from labeled data during training and dynamically updates its parameters using unlabeled test data at inference time. It unifies parameter learning, temporal evolution, prior refinement, and prediction within a single probabilistic state space architecture. This formulation enables recursive online parameter updates and principled uncertainty quantification, yielding a general and robust adaptation mechanism. The approach provides both theoretical grounding and an effective solution for online prediction under distributional shifts.
This study addresses prediction performance degradation caused by covariate shift by proposing an adaptive importance-weighted model averaging method. The approach constructs a family of estimators through exponentiated density ratios and treats the degree of weighting correction as a source of uncertainty. By optimizing convex combination weights in a data-driven manner, it effectively balances bias and variance to achieve asymptotically optimal prediction on the target domain. A key innovation lies in integrating importance weighting into the model averaging framework while establishing theoretical optimality. Empirical evaluations on both simulated and real-world datasets demonstrate that the proposed method achieves competitive predictive performance compared to existing approaches.
This work addresses the instability of Test-Time Training (TTT) under distribution shift, which stems from its sensitivity to hyperparameters and the absence of theoretical guidance. The authors reinterpret TTT through the lens of decision theory as implicit Bayesian inference under a kernel mechanism. They propose a PAC-Bayes–guaranteed method that adaptively selects the number of update steps based on prompt evidence and characterize the Bayes-optimal update subspace within a linear Gaussian correction model to inform Transformer module selection. By integrating Gaussian processes, spectral analysis, and Bayesian inference, the study establishes a theoretical framework for TTT, revealing conditions under which fixed update strategies fail and providing principled foundations for adaptive update directions and step sizes, thereby effectively mitigating TTT’s instability.
This work addresses the lack of a standardized protocol in test-time adaptation for time series forecasting and the limited robustness of existing methods under distribution shifts. The study reformulates test-time adaptation from the perspective of protocol rigor, introducing a clean adaptation protocol that relies solely on observed ground-truth values. Through frequency-domain analysis, it reveals fundamental limitations in current approaches. Building upon these insights, the authors design a lightweight frequency-domain parameterized calibration module that consistently achieves strong performance across diverse datasets, forecast horizons, and backbone models. The proposed method significantly reduces parameter count compared to existing techniques while maintaining architectural clarity and computational efficiency.