Score
Designs and implements models or procedures that predict and apply corrective residuals to the outputs of an existing ‘anchor’ model so adjusted predictions better match the target residual distribution and reduce systematic errors under train–test mismatch. This competence covers building small neural correctors or adapters, residual distribution matching and rematch procedures, freezing structural logit or coefficient components while training target-updates, combining anchor signals with target updates, and analyzing calibration, ensemble dispersion, and guarantees (for example preserving mean-residual frameworks) under distribution shift or reduced context.
This work addresses the limited predictive diversity and poor out-of-distribution (OOD) uncertainty quantification of single pretrained models under distribution shift. The authors propose a Perturb-and-Correct approach that constructs a posterior ensemble using only a single model by applying random perturbations to hidden layers of the pretrained network and subsequently correcting them via least-squares affine transformations. This method uniquely exploits the affine redundancy inherent in neural networks, enhancing OOD prediction diversity and uncertainty calibration without compromising in-distribution performance. Empirical evaluations demonstrate that the proposed technique achieves a superior or competitive trade-off between in-distribution accuracy and OOD detection on MuJoCo dynamics prediction and CIFAR-10 OOD benchmarks compared to existing posterior ensemble baselines.
本文针对训练数据中罕见大偏移问题,提出CVaR锚回归方法,通过使用尾部平均代替平方均值残差的平均,以提高对罕见偏移的鲁棒性同时保持常见环境下的准确性。
This study addresses the scarcity of defect supervision and the reliance on manual product descriptions in few-shot industrial anomaly detection by proposing a two-stage asymmetric prompt adaptation framework. The method decouples anomaly semantic acquisition from target appearance adaptation: it first learns transferable text anchors from auxiliary data, then fixes these anchors while adapting a new branch using only target normal samples to jointly represent normality. Through dual-normal-branch inheritance adaptation and text anchor regularization, the approach eliminates dependence on category-specific templates without requiring synthetic anomalies. Evaluated under 1/2/4-shot settings across the MVTec-AD and VisA datasets, the proposed method achieves competitive performance in both anomaly detection and localization.
This work addresses systematic errors in frozen pretrained models that recurrently manifest in structured scenarios and resist correction through fine-tuning. To tackle this, the authors propose CRAFTER, a framework that, while keeping the backbone frozen, uniquely focuses feature engineering on modeling the failure process itself. CRAFTER automatically discovers interpretable corrective features by mining prediction residuals and employs them to drive a lightweight post-hoc corrector. It integrates compositional input channel search with large language model–generated named features, binary flags, and executable code, followed by gated unified validation for robust feature selection. A source-agnostic evaluation pipeline precisely attributes feature contributions. Evaluated across six datasets and six frozen backbones, CRAFTER nearly doubles the correction efficacy of existing methods, reduces error by up to 27% even for the weakest backbone, and demonstrates consistent robustness across diverse LLMs and fine-tuned models.
This study addresses the scarcity of tabular data under covariate shift and the instability of conventional augmentation methods caused by fitting outdated source distributions. To this end, we propose the IGDPR framework, which pioneers the integration of invariant potential functions into the diffusion sampling process to align stable decision boundaries and generate task-relevant samples. Furthermore, prototype clustering and reweighting strategies are incorporated to filter noise and assess sample reliability, thereby mitigating overfitting. Experimental results demonstrate that the proposed framework significantly improves synthetic data quality in real-world scenarios, effectively enhancing model robustness and generalization to unseen environments.
This study addresses the lack of binding between intervention directions and named semantics in latent spaces, as well as the problem of coordinate ambiguity. To this end, it proposes an anchor observer framework grounded in known perturbation signatures. Methodologically, by integrating capacity-matched least-squares prediction, a coordinate transport algorithm, a rank-aware abstention strategy, and fault injection testing, the approach effectively decouples prediction, semantic support, and scientific validation. Synthetic vascular auditing experiments demonstrate that this mechanism achieves predictive equivalence at 1e-15 precision, successfully identifies all injected errors, and rejects unsupported interpretations. Consequently, this work establishes a rigorous, executable separation paradigm for interpretable editing.
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.
This study addresses the failure of conformal prediction coverage guarantees caused by post-deployment data distribution shifts. It proposes Exponential Tilting Reweighting Alignment (ExTRA), a framework for modeling distribution shifts that systematically compares two adaptation strategies: weight calibration and prediction tilting. The analysis reveals that while prediction tilting can reduce set sizes under specific conditions, it may compromise coverage validity. Experiments demonstrate that ExTRA decreases prediction set lengths by 30% in synthetic regression tasks; however, it induces significant coverage degradation in classification or information-deficient settings and yields no consistent benefits on real-world data. Determining the precise applicability conditions for this approach remains an open challenge.
This study addresses the unreliability of historical statistics caused by distribution shifts in continual test-time adaptation by proposing the GAIN framework. Following a "historical proposal, gain decision" principle, GAIN introduces a posterior predictive evaluator to quantify estimation uncertainty and dynamically determines sample-level intervention intensity through backpropagation-free one-dimensional optimization, thereby effectively suppressing the propagation of unreliable corrections. Additionally, it maintains compact target statistics to enable efficient online updates. Evaluated on ImageNet-C, GAIN achieves 61.9% accuracy with calibration comparable to the source model while accelerating inference by 15.9× over baselines, demonstrating high accuracy, reliable calibration, and computational efficiency simultaneously.
This work addresses the challenge of verifying covariate balance in covariate shift adaptation by proposing a sequentially valid, anytime-stoppable validation framework. Built upon time-uniform confidence sequences, the method dynamically monitors covariate balance for a pre-specified function class within a prescribed tolerance band and terminates as soon as all target moments fall within this band, thereby certifying balance. Its key contribution lies in providing, for the first time, a locally and absolutely valid certification of balance for any adjustment strategy, supporting data-dependent stopping times while rigorously controlling the probability of erroneous balance confirmation. Integrated with KL-divergence-based drift diagnostics and detection of admissible adjustment regions, experiments demonstrate the method’s superior performance in error rate control, locality with respect to function classes, effectiveness in drift diagnosis, and guaranteed coverage in conformal prediction.