residual-adapter online adaptation

Designs and implements systems that attach small residual adapter modules to pretrained models and update only those modules online using parameter-efficient fine-tuning techniques, so the forecaster adapts with low compute and memory cost. This work includes the online update algorithms (e.g., solving accepted-update convex subproblems, projected online gradient descent), update-control logic (accept/reject), and auditable telemetry to record adaptation history.

residual-adapteronlineadaptation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Parameter-Efficient Continual Fine-Tuning: A Survey

Apr 18, 2025
EN
Eric Nuertey Coleman
🏛️ University of Pisa | Politecnico di Torino | University of Auckland | Indian Institute of Technology | University of Warwick

This paper addresses the fundamental challenge of balancing catastrophic forgetting and parameter efficiency when large pre-trained models continuously adapt to dynamic task streams. To this end, we propose the first unified theoretical framework for Parameter-Efficient Continual Fine-Tuning (PECFT). Our framework systematically organizes existing approaches along three dimensions: method taxonomy, evaluation metrics, and core challenges—integrating Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., adapters, LoRA, prompt tuning) with continual learning strategies (e.g., replay, regularization, architecture expansion). Through a comprehensive review of over 100 studies, we identify key trade-offs between performance and efficiency, and pinpoint scalable memory mechanisms and task-aware parameter updates as critical research frontiers. This work bridges a significant gap at the intersection of continual learning and PEFT, providing both theoretical foundations and practical guidelines for efficient, sustainable adaptation of large language models.

Addressing catastrophic forgetting in continual learning scenariosEnhancing parameter-efficient fine-tuning for dynamic environmentsSurveying methods for lifelong adaptation of large pre-trained models

Must-Read Papers

Most classic and influential ideas
View more

Online time series prediction using feature adjustment

Sep 03, 2025
XH
Xiannan Huang
🏛️ Tongji University

To address degraded model adaptability in online multistep time-series forecasting—caused by data distribution drift and delayed ground-truth feedback—this paper proposes ADAPT-Z. Methodologically, ADAPT-Z abandons conventional parameter fine-tuning and instead models the dynamics of latent factors. It introduces an adapter module that fuses current features with historical gradient information within a learned Z-space, enabling persistent tracking and incremental self-adaptation of feature representations. This design mitigates gradient mismatch induced by label delay and enhances robustness to non-stationary data. Empirical evaluation across multiple benchmark datasets demonstrates that ADAPT-Z significantly outperforms static baselines and state-of-the-art online learning methods, achieving superior generalization and sustained adaptive capability under streaming conditions.

Addresses distribution shift in online time series forecastingProposes updating feature representations of latent factorsSolves delayed feedback issue in multi-step predictions

Lightweight Online Adaption for Time Series Foundation Model Forecasts

Feb 18, 2025
TL
Thomas L. Lee
🏛️ University of Edinburgh | Huawei

To address the challenge of online adaptation to dynamic data distributions after deployment of time-series foundation models, this paper proposes AdapTS, a lightweight online adaptation framework. Methodologically, AdapTS decouples foundation model inference from adapter learning via two novel modules: AdapTS-Forecaster—a linear or MLP-based lightweight time-series modeling component—and AdapTS-Weighter—a gradient-free, dynamically weighted fusion mechanism. It further incorporates online distribution estimation and real-time feedback-driven calibration. Crucially, AdapTS enables zero-shot fine-tuning and low-overhead adaptation without modifying the frozen foundation model. Evaluated across multiple standard benchmarks, AdapTS consistently improves the average MSE of mainstream time-series foundation models by 7.2%–15.8%, while incurring less than 3 ms additional inference latency. This demonstrates substantial gains in both predictive accuracy and practical deployability.

AdapTS mechanism improves FM performanceEnhance FM forecasts via online feedbackLightweight online adaption for time series

Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models

Nov 06, 2025
YW
Yuxiao Wang
🏛️ University of Pennsylvania | Washington University in St. Louis | Stevens Institute of Technology | University of Southern California

Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.

Addressing structural instability through graph-based dependency modelingImproving parameter efficiency and training stability via modular adaptersReducing computational costs in multi-task adaptation of large-scale models

Staying Alive: Online Neural Network Maintenance and Systemic Drift

Mar 22, 2025
JE
Joshua E. Hammond
🏛️ The University of Texas at Austin | ExxonMobil

This paper addresses the challenge of efficient online adaptation when system drift occurs post-deployment. To avoid costly full retraining or fine-tuning, we propose a dynamic weight update method grounded in Subset Extended Kalman Filtering (SEKF). SEKF dynamically identifies a critical parameter subset via loss gradient analysis and recursively updates only those weights within an Extended Kalman Filter framework, jointly optimizing accuracy and computational efficiency. The approach significantly reduces sensitivity to hyperparameter tuning. Evaluated on four dynamic regression tasks, it matches or exceeds the accuracy of fine-tuning baselines while reducing per-iteration latency by orders of magnitude. Key advantages include low computational overhead, high temporal responsiveness, and strong robustness to distributional shifts. Overall, SEKF establishes a novel paradigm for online continual learning in neural networks.

Maintaining accuracy efficiently with reduced hyperparameter tuningOnline neural network maintenance during systemic driftUpdating model weights without retraining or finetuning

Proactive Model Adaptation Against Concept Drift for Online Time Series Forecasting

Dec 11, 2024
LZ
Lifan Zhao
🏛️ Shanghai Jiao Tong University

In online time-series forecasting, concept drift—particularly temporal misalignment induced by label delay—causes models to persistently adapt to outdated concepts. To address this, we propose Proceed, a proactive model adaptation framework that departs from conventional reactive update paradigms by introducing *proactive adaptation*: it estimates the concept shift between training and current test samples to guide a generative parameter adapter in actively calibrating model parameters. Proceed comprises a generalizable drift estimation module, a lightweight parameter translation mechanism, and meta-training on synthetically generated diverse drift data to enhance robustness. Evaluated across five real-world datasets and multiple backbone forecasting models, Proceed consistently outperforms state-of-the-art online learning methods, achieving an average 12.7% reduction in MAE and significantly improving resilience to dynamic distributional shifts.

Addresses concept drift in time seriesEnhances forecast model resilienceProposes proactive model adaptation framework

Latest Papers

What's happening recently
View more

This work addresses the challenge of balancing label delay and stringent computational budgets in online time-series forecasting, where intelligent timing of model updates is crucial. The authors propose ADOWIP, a novel framework that formulates update decisions through a priority-gated mechanism driven by observed losses: updates are triggered via residual adapters only when the decision loss—revealed upon delayed feedback—exceeds an empirically calibrated quantile and sufficient budget remains. Integrating a sealed delay queue, projected online gradient descent, and conditional few-shot gating, ADOWIP delivers an auditable and feasible update policy under hard budget constraints. Evaluated on capacity planning tasks such as ETT and UCI Bike datasets, ADOWIP significantly reduces decision loss and consistently outperforms baselines on the full-year Capital Bikeshare data, with statistical significance confirmed by Holm-corrected multiple hypothesis testing.

compute budgetdecision lossdelayed feedback

This work investigates whether adaptive patching genuinely outperforms carefully tuned uniform patching in time series forecasting and under what conditions. By formulating patching as a bandwidth-constrained bitrate allocation problem, the authors derive an explicit threshold that dynamic patching must satisfy and demonstrate that local complexity alone does not guarantee the optimality of non-uniform patching. Under representation-aware optimal training, alignment gains collapse near the uniform patching baseline, suggesting that tuned uniform patching should serve as the proper benchmark for evaluating adaptive methods. Leveraging theoretical analysis based on a quadratic surrogate model and strong convexity assumptions, along with controlled experiments across three architectural variants—holding backbone, data, and training protocol fixed—the study finds that tuned uniform patching matches or exceeds adaptive approaches overall on standard long-term forecasting benchmarks, with significant gains from adaptive methods appearing only in specific method–dataset combinations.

adaptive patchingbitrate allocationforecasting loss

This work addresses the problem of efficiently recalibrating arbitrary online prediction sequences to satisfy calibration while incurring minimal excess error. The authors propose an online algorithm grounded in an enhanced simultaneous Blackwell approachability framework, which for the first time achieves the optimal $(\varepsilon, \varepsilon^2)$-recalibration rate simultaneously with calibeating for Lipschitz proper losses. This result resolves an open question regarding the near-optimal joint performance of these two desiderata and establishes the theoretical optimality of this trade-off under squared loss. The algorithm naturally extends to multi-hint settings, applies broadly to smooth proper losses, and enjoys strong theoretical guarantees. Empirical evaluations demonstrate its significant superiority over existing methods in classification tasks under distribution shift.

calibrationdistribution shiftexcess error

This work addresses the challenge of online adaptation for black-box time series foundation models when access to internal parameters is unavailable. The authors propose ORCA, a novel approach that explicitly learns the mapping between prediction errors and input-output contextual information, enabling dynamic correction of forecasts through residual modeling—without modifying the original model or relying on gradient-based updates. By circumventing the conventional paradigm of white-box fine-tuning, ORCA demonstrates superior performance over existing black-box adaptation strategies across five state-of-the-art foundation models and eight benchmark datasets, establishing its effectiveness and broad applicability.

Black-box AdaptationContext of ErrorsOnline Adaptation