test-time adaptation

Designs, builds, and evaluates methods that adapt model parameters, prompts, or outputs at inference time using incoming unlabeled inputs or feedback to mitigate distribution shift and improve performance without labeled post‑deployment data. This includes online and causal/non‑causal update protocols, self‑supervised and zeroth‑order optimization, multi‑hypothesis and negation‑aware selection strategies, feedback‑driven tuning, and analyses of stability, convergence, and robustness under query or computation constraints.

test-timeadaptation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge in performative prediction where model deployment induces distributional shifts that complicate optimization. Existing approaches often rely on strong assumptions about the loss function and data distribution, limiting their applicability. To overcome this, the paper proposes a gradient-based adaptive optimization algorithm that explicitly estimates deployment-induced distribution shifts via finite differences, thereby accommodating a broader class of losses and distributions without stringent assumptions. The method supports high-dimensional optimization and incorporates a sample-efficient approximation strategy to reduce data requirements. Theoretical analysis establishes convergence guarantees for the proposed algorithm. Empirical results demonstrate that it converges faster and more stably than existing methods, exhibiting superior robustness and practicality across diverse experimental settings.

distribution shiftgradient-based methodsloss functions

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Rethinking Distribution Shifts: Empirical Analysis and Inductive Modeling for Tabular Data

Jul 11, 2023
JL
Jiashuo Liu
🏛️ Tsinghua University | Columbia University

Distribution shift in tabular data lacks empirical grounding, and existing robust learning methods—such as distributionally robust optimization (DRO)—frequently fail in real-world settings. Prior work predominantly assumes covariate shift (X-shift), whereas empirical analysis reveals that label-conditional shift (Y|X-shift) is both more prevalent and dominant in practice. Method: The authors construct a comprehensive benchmark platform spanning five real-world tabular datasets and 60,000 experimental configurations, enabling large-scale ablation studies on robustness. Contribution/Results: (1) Theory-driven methods like DRO offer no consistent advantage over empirical risk minimization (ERM); (2) implementation details—including model selection, hyperparameter tuning, and imbalance handling—exert far greater influence on robustness than the design of ambiguity sets; (3) a data-driven, inductive paradigm should supplant prior structural assumptions. These findings challenge the conventional paradigm in distributionally robust learning and provide reproducible, empirically grounded guidance for co-optimizing algorithms and data.

Analyzing real-world distribution shifts in tabular datasetsEvaluating robust algorithms' performance against empirical shiftsIdentifying implementation factors affecting distributionally robust optimization

This work addresses the limitation of existing static tabular datasets, which lack temporal structure and thus hinder the evaluation of model adaptability under controlled distribution shifts. To overcome this, the authors propose a clustering-based framework that transforms static data into controllable, evolving data streams through cluster-based partitioning and structured perturbations. Integrating the ADWIN drift detector with a sliding-window retraining mechanism, the framework systematically evaluates adaptation strategies across six model families, including tree ensembles and online learners. Experiments on five benchmark datasets for classification and regression demonstrate that the proposed methods—particularly Clustered Local ADWIN—accurately model and efficiently respond to localized drifts in feature space, significantly outperforming baseline approaches.

concept driftdistribution shiftmodel adaptation

Hyperparameter Tuning and Model Evaluation in Causal Effect Estimation

Mar 02, 2023
DM
Damian Machlanski
🏛️ University of Essex

In causal effect estimation, the absence of standardized hyperparameter tuning evaluation criteria impedes reliable model selection and creates a substantial gap between commonly used metrics and true performance. This paper systematically investigates the interplay between hyperparameter tuning and evaluation, jointly analyzing estimators (T-/X-/R-Learner), base learners (random forests, gradient boosting, neural networks), and evaluation metrics (IPW, DR, PEHE) across four benchmark datasets. Key findings are: (1) thorough hyperparameter tuning eliminates performance differences among mainstream causal estimators; (2) the choice of evaluation strategy exerts greater influence on final performance than either the estimator type or base learner architecture; and (3) existing evaluation metrics underestimate the performance gain from optimal model selection by over 35% on average. These results demonstrate that hyperparameter tuning is the primary determinant of causal estimation accuracy, underscoring an urgent need for more robust, theoretically grounded evaluation paradigms in causal machine learning.

Complex model selection involving multiple components complicates causal inferenceLack of consensus on tuning metrics for causal effect estimation modelsNo ideal metric exists for hyperparameter tuning of causal estimators

Latest Papers

What's happening recently
View more

This study addresses the limited statistical gains achievable by existing data-driven optimization methods in the absence of informative priors. It provides a systematic analysis of the “directional perturbation” empirical optimization (EO+) framework, establishing for the first time a unified theoretical perspective that demonstrates only second-order improvements are attainable without geometrically valid side information—thereby confirming a “no free lunch” negative result. The work breaks new ground by showing that incorporating geometrically effective side information and substantially increasing a key hyperparameter enables first-order statistical improvement. Furthermore, it introduces a gain-maximization mechanism based on excess risk estimation, significantly enhancing decision efficiency. By integrating bootstrap resampling, control variates, distributionally robust optimization, and transfer learning, this research bridges data-driven optimization with Monte Carlo variance reduction theory.

data-driven optimizationempirical optimizationfirst-order improvement

This work addresses the challenge of balancing label delay and stringent computational budgets in online time-series forecasting, where intelligent timing of model updates is crucial. The authors propose ADOWIP, a novel framework that formulates update decisions through a priority-gated mechanism driven by observed losses: updates are triggered via residual adapters only when the decision loss—revealed upon delayed feedback—exceeds an empirically calibrated quantile and sufficient budget remains. Integrating a sealed delay queue, projected online gradient descent, and conditional few-shot gating, ADOWIP delivers an auditable and feasible update policy under hard budget constraints. Evaluated on capacity planning tasks such as ETT and UCI Bike datasets, ADOWIP significantly reduces decision loss and consistently outperforms baselines on the full-year Capital Bikeshare data, with statistical significance confirmed by Holm-corrected multiple hypothesis testing.

compute budgetdecision lossdelayed feedback

This work addresses the challenge of enabling predictive systems to dynamically adapt their behavior based on contextual information for personalized inference. To this end, it proposes a unified framework that maps context into adaptation parameters for prediction and, for the first time, establishes a mathematical equivalence between explicit parameter adaptation and implicit expert routing under kernel ridge regression. The framework theoretically unifies diverse methodologies—including varying-coefficient models, local regression, prompt engineering, retrieval-augmented approaches, and mixture-of-experts—under fixed features and squared loss. Key contributions include deriving a general formulation for context-adaptive inference, proposing practical design principles and evaluation metrics such as adaptation efficiency and routing stability, and highlighting critical open problems concerning identifiability and robustness under distributional shifts.

context-adaptive inferencedistribution shiftfoundation models

Hot Scholars

HT

Hanghang Tong

University of Illinois at Urbana-Champaign
Large Scale Data MiningGraph MiningSocial NetworksHealthcare
XZ

Xiatian Zhu

University of Surrey
Machine LearningComputer Vision
LZ

Lihua Zhou

Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, CAS
Machine LearningTransfer Learning
JH

Jingrui He

University of Illinois at Urbana-Champaign
Machine LearningData MiningSocial NetworksMedical Informatics
EG

Eric Granger

Professor of Systems Engineering, École de technologie supérieure, LIVIA, ILLS, REPARTI
Machine LearningComputer VisionPattern RecognitionAffective Computing