Score
Designs and implements systems that attach small residual adapter modules to pretrained models and update only those modules online using parameter-efficient fine-tuning techniques, so the forecaster adapts with low compute and memory cost. This work includes the online update algorithms (e.g., solving accepted-update convex subproblems, projected online gradient descent), update-control logic (accept/reject), and auditable telemetry to record adaptation history.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
This paper addresses the fundamental challenge of balancing catastrophic forgetting and parameter efficiency when large pre-trained models continuously adapt to dynamic task streams. To this end, we propose the first unified theoretical framework for Parameter-Efficient Continual Fine-Tuning (PECFT). Our framework systematically organizes existing approaches along three dimensions: method taxonomy, evaluation metrics, and core challenges—integrating Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., adapters, LoRA, prompt tuning) with continual learning strategies (e.g., replay, regularization, architecture expansion). Through a comprehensive review of over 100 studies, we identify key trade-offs between performance and efficiency, and pinpoint scalable memory mechanisms and task-aware parameter updates as critical research frontiers. This work bridges a significant gap at the intersection of continual learning and PEFT, providing both theoretical foundations and practical guidelines for efficient, sustainable adaptation of large language models.
To address degraded model adaptability in online multistep time-series forecasting—caused by data distribution drift and delayed ground-truth feedback—this paper proposes ADAPT-Z. Methodologically, ADAPT-Z abandons conventional parameter fine-tuning and instead models the dynamics of latent factors. It introduces an adapter module that fuses current features with historical gradient information within a learned Z-space, enabling persistent tracking and incremental self-adaptation of feature representations. This design mitigates gradient mismatch induced by label delay and enhances robustness to non-stationary data. Empirical evaluation across multiple benchmark datasets demonstrates that ADAPT-Z significantly outperforms static baselines and state-of-the-art online learning methods, achieving superior generalization and sustained adaptive capability under streaming conditions.
To address the challenge of online adaptation to dynamic data distributions after deployment of time-series foundation models, this paper proposes AdapTS, a lightweight online adaptation framework. Methodologically, AdapTS decouples foundation model inference from adapter learning via two novel modules: AdapTS-Forecaster—a linear or MLP-based lightweight time-series modeling component—and AdapTS-Weighter—a gradient-free, dynamically weighted fusion mechanism. It further incorporates online distribution estimation and real-time feedback-driven calibration. Crucially, AdapTS enables zero-shot fine-tuning and low-overhead adaptation without modifying the frozen foundation model. Evaluated across multiple standard benchmarks, AdapTS consistently improves the average MSE of mainstream time-series foundation models by 7.2%–15.8%, while incurring less than 3 ms additional inference latency. This demonstrates substantial gains in both predictive accuracy and practical deployability.
Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.
This paper addresses the challenge of efficient online adaptation when system drift occurs post-deployment. To avoid costly full retraining or fine-tuning, we propose a dynamic weight update method grounded in Subset Extended Kalman Filtering (SEKF). SEKF dynamically identifies a critical parameter subset via loss gradient analysis and recursively updates only those weights within an Extended Kalman Filter framework, jointly optimizing accuracy and computational efficiency. The approach significantly reduces sensitivity to hyperparameter tuning. Evaluated on four dynamic regression tasks, it matches or exceeds the accuracy of fine-tuning baselines while reducing per-iteration latency by orders of magnitude. Key advantages include low computational overhead, high temporal responsiveness, and strong robustness to distributional shifts. Overall, SEKF establishes a novel paradigm for online continual learning in neural networks.
In online time-series forecasting, concept drift—particularly temporal misalignment induced by label delay—causes models to persistently adapt to outdated concepts. To address this, we propose Proceed, a proactive model adaptation framework that departs from conventional reactive update paradigms by introducing *proactive adaptation*: it estimates the concept shift between training and current test samples to guide a generative parameter adapter in actively calibrating model parameters. Proceed comprises a generalizable drift estimation module, a lightweight parameter translation mechanism, and meta-training on synthetically generated diverse drift data to enhance robustness. Evaluated across five real-world datasets and multiple backbone forecasting models, Proceed consistently outperforms state-of-the-art online learning methods, achieving an average 12.7% reduction in MAE and significantly improving resilience to dynamic distributional shifts.
This work addresses the challenge of balancing label delay and stringent computational budgets in online time-series forecasting, where intelligent timing of model updates is crucial. The authors propose ADOWIP, a novel framework that formulates update decisions through a priority-gated mechanism driven by observed losses: updates are triggered via residual adapters only when the decision loss—revealed upon delayed feedback—exceeds an empirically calibrated quantile and sufficient budget remains. Integrating a sealed delay queue, projected online gradient descent, and conditional few-shot gating, ADOWIP delivers an auditable and feasible update policy under hard budget constraints. Evaluated on capacity planning tasks such as ETT and UCI Bike datasets, ADOWIP significantly reduces decision loss and consistently outperforms baselines on the full-year Capital Bikeshare data, with statistical significance confirmed by Holm-corrected multiple hypothesis testing.
This work investigates whether adaptive patching genuinely outperforms carefully tuned uniform patching in time series forecasting and under what conditions. By formulating patching as a bandwidth-constrained bitrate allocation problem, the authors derive an explicit threshold that dynamic patching must satisfy and demonstrate that local complexity alone does not guarantee the optimality of non-uniform patching. Under representation-aware optimal training, alignment gains collapse near the uniform patching baseline, suggesting that tuned uniform patching should serve as the proper benchmark for evaluating adaptive methods. Leveraging theoretical analysis based on a quadratic surrogate model and strong convexity assumptions, along with controlled experiments across three architectural variants—holding backbone, data, and training protocol fixed—the study finds that tuned uniform patching matches or exceeds adaptive approaches overall on standard long-term forecasting benchmarks, with significant gains from adaptive methods appearing only in specific method–dataset combinations.
This work addresses the problem of efficiently recalibrating arbitrary online prediction sequences to satisfy calibration while incurring minimal excess error. The authors propose an online algorithm grounded in an enhanced simultaneous Blackwell approachability framework, which for the first time achieves the optimal $(\varepsilon, \varepsilon^2)$-recalibration rate simultaneously with calibeating for Lipschitz proper losses. This result resolves an open question regarding the near-optimal joint performance of these two desiderata and establishes the theoretical optimality of this trade-off under squared loss. The algorithm naturally extends to multi-hint settings, applies broadly to smooth proper losses, and enjoys strong theoretical guarantees. Empirical evaluations demonstrate its significant superiority over existing methods in classification tasks under distribution shift.
This work addresses the challenge of online adaptation for black-box time series foundation models when access to internal parameters is unavailable. The authors propose ORCA, a novel approach that explicitly learns the mapping between prediction errors and input-output contextual information, enabling dynamic correction of forecasts through residual modeling—without modifying the original model or relying on gradient-based updates. By circumventing the conventional paradigm of white-box fine-tuning, ORCA demonstrates superior performance over existing black-box adaptation strategies across five state-of-the-art foundation models and eight benchmark datasets, establishing its effectiveness and broad applicability.