Score
Design and apply analytical frameworks and evaluation protocols that decompose time-varying predictive error into separate contributions from model-change (updates, learning dynamics, or optimization) and environment-change (distributional or semantic shifts). Use these decompositions to derive theoretical comparisons of optimization procedures (e.g., gradient-descent vs. reinforcement-learning style updates) and to produce bounds or metrics for future-domain generalization and semantic out-of-distribution rejection.
This work addresses the challenge of data distribution shifts arising jointly from endogenous factors—induced by model decisions—and exogenous factors—stemming from spontaneous environmental changes—after deployment of predictive models. To tackle this issue, the paper introduces a partially actionable prediction framework that, for the first time, unifies the modeling of both types of distributional shifts and formalizes the notions of online actionable stability and optimality. Extending classical actionable prediction theory, the proposed framework establishes a dynamic online learning paradigm that integrates distribution evolution modeling with optimization analysis. Theoretical results demonstrate that, under reasonable assumptions, online strategies such as repeated retraining can effectively track distributional changes and achieve stable convergence.
Existing evaluation methods struggle to disentangle whether performance degradation under temporal distribution shifts stems from insufficient model adaptability or increased data difficulty. To address this, this work introduces a novel approach that decouples model adaptability from the inherent difficulty of temporal data for the first time. The authors propose three dynamic metrics based on performance trajectories, which capture the adaptation process through dynamic evaluation and comparative analysis. These new metrics uncover fine-grained adaptation patterns obscured by conventional assessment techniques, substantially enhancing the interpretability and depth of understanding of temporal robustness in machine learning models.
This work addresses the challenge in performative prediction where model deployment induces distributional shifts that complicate optimization. Existing approaches often rely on strong assumptions about the loss function and data distribution, limiting their applicability. To overcome this, the paper proposes a gradient-based adaptive optimization algorithm that explicitly estimates deployment-induced distribution shifts via finite differences, thereby accommodating a broader class of losses and distributions without stringent assumptions. The method supports high-dimensional optimization and incorporates a sample-efficient approximation strategy to reduce data requirements. Theoretical analysis establishes convergence guarantees for the proposed algorithm. Empirical results demonstrate that it converges faster and more stably than existing methods, exhibiting superior robustness and practicality across diverse experimental settings.
This work addresses the significant performance degradation of reinforcement learning (RL) agents under distribution shifts arising from mismatches or dynamics between training and deployment environments, a challenge exacerbated by the lack of a systematic understanding of their causal origins. By modeling agent–environment interaction through the lens of partially observable Markov decision processes (POMDPs), the study decomposes the RL framework into causal components—states, observations, policies, rewards, and transitions—and, incorporating temporal boundaries of shift occurrence, offers the first unified characterization of distribution shifts grounded in causal mechanisms. It distinguishes between internal (agent-driven) and external (environment-driven) sources and introduces a novel taxonomy encompassing explicit, implicit, and hybrid shifts. This framework establishes a structured classification and evaluation system for distribution shifts, enabling systematic analysis and targeted improvements of RL robustness.
This work addresses the poor out-of-distribution (OOD) generalization, lack of theoretical guarantees, and limited interpretability of recurrent neural networks (RNNs) on temporal data. By modeling the post-training RNN state dynamics as a nonlinear closed-loop system, the authors introduce Koopman operator theory—applied here for the first time to RNNs—to approximate this system with a linear representation. Combining this linearization with spectral analysis, they rigorously quantify the worst-case impact of domain shift on generalization error. Based on this analysis, they derive a generalization error bound for non-i.i.d. temporal data and propose an interpretable, robust domain generalization training method. Experiments across multiple temporal tasks demonstrate that the proposed approach significantly reduces OOD generalization error and enhances model robustness to domain shifts.
This work addresses the high cost of large model fine-tuning by tackling the challenge of accurately predicting post-fine-tuning performance beforehand—a task whose theoretical limits remain unclear. We formulate pre-fine-tuning performance prediction as a stochastic estimation problem under information constraints and introduce a predictive risk decomposition framework that separates it into an irreducible intrinsic limit and an optimizable variance term, thereby revealing fundamental bounds on predictability. Leveraging information theory and statistical learning theory, we establish a theoretical lower bound on variance decay through optimization and construct a predictability phase diagram that delineates three distinct task regimes. Experiments on both synthetic and real-world benchmarks validate the efficacy of this phase diagram, and our proposed budget-optimal probing strategy significantly enhances prediction efficiency, offering both theoretical grounding and practical tools for pre-fine-tuning decision-making.
This work addresses the challenge of continual adaptation in dynamic open-world settings, where models must generalize under covariate shift while reliably rejecting semantically out-of-distribution (OOD) inputs. Existing approaches lack explicit modeling of future environmental shifts. To bridge this gap, the paper establishes the first theoretical framework for dynamic OOD detection and introduces a reinforcement learning (RL)-guided optimizer that augments standard gradient descent with an RL-based correction term, explicitly minimizing long-term semantic OOD false positive rates. By decomposing generalization error into components attributable to model evolution and environmental dynamics over time, the method directly optimizes future generalization capability. Empirical results demonstrate that the proposed optimizer consistently outperforms conventional counterparts in both future domain generalization and semantic OOD rejection, thereby validating the efficacy of the theoretical framework.
This work addresses the instability of Test-Time Training (TTT) under distribution shift, which stems from its sensitivity to hyperparameters and the absence of theoretical guidance. The authors reinterpret TTT through the lens of decision theory as implicit Bayesian inference under a kernel mechanism. They propose a PAC-Bayes–guaranteed method that adaptively selects the number of update steps based on prompt evidence and characterize the Bayes-optimal update subspace within a linear Gaussian correction model to inform Transformer module selection. By integrating Gaussian processes, spectral analysis, and Bayesian inference, the study establishes a theoretical framework for TTT, revealing conditions under which fixed update strategies fail and providing principled foundations for adaptive update directions and step sizes, thereby effectively mitigating TTT’s instability.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work addresses the challenge of robust modeling and prediction for nonlinear dynamical systems under distribution shift by proposing a Bayesian meta-learning framework. The approach approximates nonlinear dynamics through a linear latent space and, for the first time, employs a Matrix Normal-Inverse Wishart prior to model the Koopman operator, enabling joint quantification of epistemic and aleatoric uncertainties with closed-form posterior updates. Evaluated on real-world heavy-duty truck data and diverse simulation tasks, the method significantly outperforms existing approaches in multi-step prediction accuracy, uncertainty calibration, and robustness to distribution shifts. Furthermore, it successfully enables feasible motion planning under extreme operating conditions.